Your question is Design a Latency Aware Reasoning Agent. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are building an agentic assistant that can answer simple requests quickly but may also use iterative reasoning and tools for harder tasks. Some users care most about speed, while others care more about answer quality on complex problems.
How do you balance inference latency with the computational cost of iterative reasoning steps?