Your question is Benchmarking LLM Evaluation Systems. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
How do you evaluate and benchmark LLM evaluation systems using both automated judges and human annotation frameworks?
Explain how you would design the evaluation dataset, define quality dimensions, measure judge validity and reliability, compare automated judgments with human labels, and analyze disagreements. Address calibration, annotator disagreement, judge bias, contamination, reproducibility, cost, and how benchmark results should guide evaluator improvements.