Your question is Optimizing Latency-Sensitive Inference. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
Explain how you would optimize a latency-sensitive AI inference pipeline.
Discuss how you would measure latency, identify bottlenecks, and choose among model, hardware, serving, and architectural optimizations. Address online versus batch work, AWS implementation options, cost, quality regressions, and production failure handling.