Your question is Compare Model Performance Across Datasets. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are comparing the same model on multiple datasets, and the metrics do not line up the way you expected. One dataset looks strong, another is noticeably weaker, and a third sits in between. The team wants a clear read on whether the model is stable across data sources or if the results point to a real generalization problem.
How do you evaluate the performance of a machine learning model across different datasets?