Your question is LLM Evaluation Metrics. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
How do you approach LLM evaluation? What metrics do you prioritize for a summarization task versus a Q&A task?
Explain how you would build an evaluation strategy rather than relying on a single score. Cover reference-based metrics, rubric-based human or LLM evaluation, hallucination and faithfulness checks, refusal behavior, latency, cost, and production monitoring. Distinguish metrics that measure summary quality from those that measure answer correctness and grounding.
Provide a practical evaluation design, including representative test data, error analysis, and release criteria.