Your question is Build Reliable Model Evaluation Process. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You have trained and shipped a machine learning model, and the team wants confidence that its performance will hold up outside the initial offline results. You need a clear evaluation process that catches overfitting, unstable thresholds, and score quality issues before the model affects users.
How do you ensure that your machine learning models are robust and reliable?