Your question is Evaluate an LLM System. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are working on an LLM-powered product feature and need a clear way to judge whether the model is good enough to ship and improve over time. The outputs are open-ended, so simple accuracy is not enough.
How do you evaluate the performance of a generative model?