Problem
Scenario
You are preparing a supervised learning dataset and notice that some fields are missing, inconsistent, or clearly noisy. You want a clean training pipeline that improves model quality without introducing leakage.
Question
How would you handle missing or noisy data in a machine learning dataset?
Example Dataset
Size·120K rows, 38 featuresTarget·Binary conversion within 14 daysFeature mix·Numerical, categorical, behavioral aggregatesData quality issues·5% to 18% missingness, outliers, inconsistent logged values
You are practicing as a guest. Sign up free to get your answer graded with AI feedback. Your draft stays right here.



