Your question is Optimizing PySpark Jobs. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
Describe how you optimize PySpark code for large-scale data processing.
Cover the changes you would make in PySpark code and Spark configuration, how you would identify bottlenecks, and the trade-offs involved. Keep the answer focused on execution plans, shuffles, partitions, joins, serialization, and memory usage.