Your question is Evaluating Agentic Workflows. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
How do you design an evaluation framework to measure the accuracy and reliability of an agentic workflow?
Describe a practical approach suitable for an agent that may plan, call tools, retrieve information, generate intermediate outputs, and complete multi-step tasks. Explain how you would define success, build representative evaluation data, select metrics, handle non-deterministic behavior, perform error analysis, and monitor quality after deployment.
Your answer should distinguish component-level evaluation from end-to-end evaluation and explain how you would validate both task outcomes and workflow reliability.