Your question is Handling Pipeline Errors at Scale. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are responsible for an automated workflow that moves and transforms operational data across internal systems. The workflow normally runs without much manual intervention, but you want a clear approach for when failures start happening repeatedly and affect a large number of records or downstream users.
What would you do if an automated workflow started creating errors at scale?