Your question is Evaluate Non Deterministic AI Agents. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are evaluating an AI agent that can take multiple valid paths to complete the same task, and its outputs vary across runs. The team wants a framework that can judge quality fairly, compare versions, and catch regressions without assuming deterministic behavior.
How do you design an evaluation framework for non-deterministic AI agents?