Problem
Write a Python or PySpark script to read a large, nested JSON dataset from cloud storage, flatten the schema, perform an aggregation, and write the output back as partitioned Parquet files.
Practicing as: Data Engineer interview at Virtual VocationsHi, I'll play your Virtual Vocations interviewer for the Data Engineer role. Answer the question above like we're in the room, and I'll respond the way a real interviewer would.
You are practicing as a guest. Sign up free to get your answer graded with AI feedback. Your draft stays right here.

