Your question is Reduce MTTR for Production Incidents. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are responsible for improving operational execution for a platform team that supports critical production services. Recent incidents have been resolved eventually, but recovery has been too slow and leadership wants a clearer approach to reducing Mean Time to Repair without increasing change risk.
What strategies do you use to reduce MTTR?