Your question is Evaluating Agentic Model Quality. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are evaluating a model that can plan, call tools, and recover from intermediate mistakes. Simple task accuracy is useful, but it misses whether the model behaves like a strong agent across multi-step work. You want a framework that captures planning quality, tool use, recovery, and consistency.
What metrics do you use to evaluate the agentic quality of a model beyond simple accuracy?