Your question is Optimizing Inference Latency. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
How do you optimize latency for inference in a distributed system for Flipkart-scale traffic?
Explain how you would reason about the serving architecture, identify latency bottlenecks, and choose optimizations across networking, feature retrieval, model execution, batching, caching, and hardware. Cover how you would validate improvements and handle tail latency, failures, model updates, feature drift, and training-serving skew.