Your question is Distributed Failures in Production. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
How would you handle distributed system failures in a production environment? Address detection, incident coordination, containment, graceful degradation, recovery, data consistency, communication, and post-incident improvements. Explain how you would prioritize actions when the failure scope and root cause are initially unclear.