Welcome to your interview.
The question is on your right: Design a Highly Available ML Serving Platform. Take a moment with it first.
Talk your thinking through with me if you like - when you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes). Discussion and graded submissions share your five interviewer interactions, so spend them well.
You are building an ML powered decision service that runs inside a distributed application. Predictions must stay available during traffic spikes, partial outages, and model rollouts, while keeping latency low and outputs consistent across regions.
How do you ensure high availability and scalability in a distributed system?