Your question is Evaluating LLMs in Production. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
How do you evaluate the performance of different LLMs in a production setting?
Explain how you would design a practical evaluation framework for comparing models after deployment. Cover offline benchmarks, task-specific quality metrics, human or expert review, production monitoring, latency, cost, reliability, safety, and handling non-deterministic outputs. Describe how you would validate that an apparent improvement is real and decide whether to promote one model over another.