Your question is Design a Highly Available ML Serving Platform. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are building an ML powered decision service that runs inside a distributed application. Predictions must stay available during traffic spikes, partial outages, and model rollouts, while keeping latency low and outputs consistent across regions.
How do you ensure high availability and scalability in a distributed system?