Your question is Data Cleaning and Modeling Pipeline. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
Given a dataset split across multiple CSV files, how would you clean the data, preprocess features with scikit-learn, and train a classifier to predict a target class?