Your question is Resolve a Production GPU Outage. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are on call for a production platform that serves GPU-backed workloads. An incident causes customer jobs to fail shortly after a deployment, and internal teams are asking for updates while service health is degrading. You need to balance fast mitigation, clear communication, and enough technical investigation to avoid making the situation worse.
Describe a time you worked on a production issue and how you resolved it. Walk through how you diagnosed the problem, decided whether to roll back or fix forward, communicated with stakeholders during the incident, and what you changed afterward to prevent recurrence.