Your question is Evaluate Retrieval Without Generation. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are building a retrieval-augmented assistant for an internal knowledge base. The team wants to know whether bad answers come from retrieval or from the model's generation step.
Walk me through how you would evaluate retrieval quality independently from generation quality.