Welcome to your interview.
The question is on your right: Handle PySpark Data Skew. Take a moment with it first.
Talk your thinking through with me if you like - when you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes). Discussion and graded submissions share your five interviewer interactions, so spend them well.
You're working on a PySpark pipeline and notice that some stages run much slower than others because a small number of keys dominate the data. You want to reason about how to diagnose and reduce skew before it causes repeated job delays.
How do you handle data skewness in a PySpark environment?