Your question is Design a Low Latency Inference Platform. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are building an online prediction service for a product that depends on model outputs in the request path. Traffic is bursty, predictions are frequent, and users notice delays immediately. The platform needs to support multiple models and evolving features without slowing down the product.
How would you architect a system to handle high-frequency model inference requests with minimal latency?