Your question is Testing Model Quality Improvement. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You retrained a model and saw better offline quality on a held-out evaluation set. Before shipping it, you want to know whether the observed lift is real or could plausibly be noise from the sample you evaluated on.
How would you determine whether an observed improvement in model quality is statistically significant?