Your question is Fault Tolerance in Data Pipelines. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You're planning a new data pipeline and want to think through how it behaves when parts of the system fail. The goal is to avoid duplicate data, silent data loss, and long recovery times when jobs, workers, or upstream sources break.
How would you implement fault tolerance in a distributed system?