Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started

Triage a Critical Service Outage

Hard
HardSecurity & InfrastructureInfrastructureToolsQualityAsked 1 times

Scenario

Scenario

You are responsible for a critical university service that authenticates users and serves protected records through an internal API. The service is now unavailable, and you have just been paged because requests are timing out across multiple clients. You also see signs that a recent configuration change may have affected both connectivity and access control.

Question

How would you approach restoring service while making sure you do not widen the blast radius or bypass security controls? What would you check first, how would you decide whether to fail closed or degrade gracefully, and how would you verify the system is safe to bring back up?

What Matters

  • Service availability for authenticated users
  • Protected institutional records and authorization boundaries
  • Internal connectivity, dependency health, and rollback safety
  • Auditability of every emergency action

What You Observe

  • Timeouts from multiple clients
  • Possible recent config or policy change
  • Unclear whether the root cause is network, identity, or dependency related
Practicing as: Software Engineer interview at Northwestern University

Hi, I'll play your Northwestern University interviewer for the Software Engineer role. Candidates describe these interviews as mostly positive and moderately difficult, so expect me to be friendly and conversational. Take your time with the question above and answer like we're in the room.

Take this as a live interview session →

You are practicing as a guest. Sign up free to get your answer graded with AI feedback. Your draft stays right here.

Sign up freeI have an account
Sign up to unlock solutions
Northwestern University Software Engineer Interview QuestionsNorthwestern University Interview Questions
Next questions
ETriage a Critical Production OutageHardScienceLogicTriage a Suspected Service CompromiseMediumCohesityTriage a Downed Production ServerMedium