Your question is Evaluate Precision and Recall Tradeoffs. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are reviewing a binary classifier that flags items for human review. The team says the model looks good overall, but reviewers are missing too many true positives and also spending time on false alarms. You need to judge how precision and recall should be evaluated for this system.
How would you evaluate precision and recall for an AI system?