Your question is Evaluate an Agentic Workflow. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are preparing to ship an LLM agent that can plan, call tools, and return final answers to users. Before launch, you need a clear way to measure whether the workflow is accurate, safe, and fast enough for production use.
How would you design an evaluation framework to measure the accuracy, safety, and latency of an agentic workflow before deploying it to production?