Your question is Evaluating Generative Output Quality. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
How do you evaluate the quality and diversity of outputs from a generative model?
Explain how you would design an evaluation framework for open-ended outputs. Address automatic metrics, human evaluation, diversity and coverage measurements, reference-based and reference-free assessment, and validation of metric reliability. Discuss how you would handle stochastic generation, task-specific quality criteria, and disagreement between metrics.