Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started

Flatten Nested JSON to Parquet

Hard
HardCodingpysparkparquetpythonAsked 1 times

Problem

Write a Python or PySpark script to read a large, nested JSON dataset from cloud storage, flatten the schema, perform an aggregation, and write the output back as partitioned Parquet files.

Practicing as: Data Engineer interview at Virtual Vocations

Hi, I'll play your Virtual Vocations interviewer for the Data Engineer role. Answer the question above like we're in the room, and I'll respond the way a real interviewer would.

Take this as a live interview session →

You are practicing as a guest. Sign up free to get your answer graded with AI feedback. Your draft stays right here.

Sign up freeI have an account
Sign up to unlock solutions
Virtual Vocations Data Engineer Interview QuestionsVirtual Vocations Interview Questions
Next questions
Flatten Nested JSON to ParquetMediumOak St. HealthFlatten Nested JSON to ParquetMediumEXL Service PhilippinesExplode Nested JSON in PySparkMedium