Your question is Design a Fault-Tolerant Web ML Stack. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are building a web application that relies on machine learning for core user-facing decisions. The product must stay usable during infrastructure failures, degraded dependencies, and bad model rollouts, while still serving predictions with acceptable quality.
What approaches do you take to ensure high availability and fault tolerance in a web application?