Your question is Compare Classification Metrics. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are reviewing a binary classifier that flags cases for manual review in a healthcare workflow. The team has four metrics on the same validation set, and they want a clear comparison of what each one says about model quality.
How would you compare model performance using precision, recall, F1-score, and ROC-AUC?