Problem
Scenario
You are working on a supervised learning problem and find that parts of the dataset are incomplete, some labels or features look noisy, and the sample may not represent the population you care about.
Question
How would you handle missing, noisy, or biased data in your research?
Example Dataset
Size·120K customer sessions, 38 featuresTarget·Binary conversion after recommendationFeatures·Numeric behavior signals, categorical profile fields, channel metadata, historical aggregatesMissing data·5% to 35% missing by feature, plus weak labels and underrepresented channelsClass balance·18% positive
Practicing as: Data Scientist interview at L'OréalHi, I'll play your L'Oréal interviewer for the Data Scientist role. Answer the question above like we're in the room, and I'll respond the way a real interviewer would.
You are practicing as a guest. Sign up free to get your answer graded with AI feedback. Your draft stays right here.


