Your question is Preprocessing Noisy Image Data. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
Given a large set of images with varying quality, how would you preprocess the data to train a robust model?
Explain a practical batch pipeline for ingestion, validation, quality assessment, normalization, augmentation, deduplication, and dataset versioning. Address corrupt files, inconsistent formats and resolutions, class imbalance, data leakage, reproducibility, and rejected-image handling. Describe how you would orchestrate the workflow, monitor data quality, and make the output available for model training.