Your question is Upgrade Databricks Production Observability. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
Databricks runs a customer-facing data platform workflow that provisions jobs, serves SQL Warehouse traffic, and powers internal on-call operations. Over the last quarter, the team has had 6 production incidents where alerts fired late, dashboards lacked enough context to isolate the fault, and mean time to resolution averaged 78 minutes. You are the DevOps Engineer responsible for leading an observability improvement project across one critical production surface in 8 weeks.
The working team includes 6 engineers: 2 platform engineers, 2 SREs, 1 software engineer from the service team, and you as the execution lead. The urgency is high because the VP of Engineering wants the new monitoring baseline in place before a large enterprise customer launch next quarter.
The Engineering Director wants faster incident detection without increasing pager fatigue. The product team wants no customer-visible downtime during rollout. Security wants auditability for alert changes and access controls. Finance has capped incremental tooling spend, while the on-call team wants fewer noisy alerts and better runbooks.