Your question is Evaluating LLM Success Beyond Perplexity. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
Describe the challenges of LLM evaluation and how you measure success beyond simple perplexity.
Discuss how you would evaluate an LLM-powered application in practice, including quality, factuality, safety, robustness, latency, and cost. Explain how you would build offline test sets, use human or LLM-based judges, monitor production behavior, and handle cases where ground-truth answers are unavailable.
Give concrete metrics, test designs, and examples of how you would use evaluation results to choose between prompts, models, retrieval strategies, or fine-tuning.