Your question is Build Reliable Model Evaluation. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You have trained and shipped a model, and the team wants confidence that its performance will hold up as usage grows and data changes over time. You need an evaluation approach that covers validation stability, score calibration, and decision threshold quality.
How do you ensure your models are robust, scalable, and accurate?