Your question is Triage a Critical Production Outage. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are responsible for a critical production system that supports internal operations and customer-facing workflows. You have just been paged because the service is unavailable, downstream requests are timing out, and the on-call dashboard shows the failure is spreading across dependent components. You do not yet know whether this is a software defect, infrastructure issue, or security event.
How would you approach this situation from the moment you are paged through recovery? What would you do to restore service safely, determine the root cause, and keep stakeholders informed while the system is down?