Your question is Observability at Fleet Scale. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are responsible for improving observability for a large distributed infrastructure footprint that now spans thousands of nodes. The current setup has basic dashboards and alerts, but signal quality, cost, and operational noise become much harder as the fleet grows.
How do you approach observability and monitoring when you have thousands of nodes?