Your question is Handle PySpark Data Skew. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You're working on a PySpark pipeline and notice that some stages run much slower than others because a small number of keys dominate the data. You want to reason about how to diagnose and reduce skew before it causes repeated job delays.
How do you handle data skewness in a PySpark environment?