Welcome to your interview.
The question is on your right: Minimizing Spark Data Shuffling. Take a moment with it first.
Talk your thinking through with me if you like - when you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes). Discussion and graded submissions share your five interviewer interactions, so spend them well.
What is data shuffling in Spark, why is it expensive, and how can you minimize or avoid it?