Your question is Handling Missing Data in Pipelines. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You're building a training data pipeline and need a consistent way to deal with incomplete records before they reach downstream models. Some fields are optional, some are critical, and missing values can come from source gaps, late data, or parsing failures.
How would you handle missing data in a dataset?