Welcome to your interview.
The question is on your right: Evaluating Without Ground Truth. Take a moment with it first.
Talk your thinking through with me if you like - when you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes). Discussion and graded submissions share your five interviewer interactions, so spend them well.
How would you evaluate the performance and accuracy of an AI agent when there is no clear ground truth dataset available, such as for Mirakl marketplace agent behaviors?