Your question is Validate Real-World Model Performance. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You have an offline model that looks strong in validation, but the team is asking whether it actually works in the field. The same score threshold is being used in production, and stakeholders want evidence that the model's decisions hold up once real users, real delays, and real labels are involved.
How do you validate real-world performance of a model beyond offline metrics?