Welcome to your interview.
The question is on your right: Partitioning vs Bucketing at Scale. Take a moment with it first.
Talk your thinking through with me if you like - when you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes). Discussion and graded submissions share your five interviewer interactions, so spend them well.
At Meta, the Ads Insights team stores large fact tables in a Hive-compatible lake and processes them with Apache Spark. A daily reporting pipeline for impression, click, and conversion events has become slow and expensive because tables are laid out inconsistently across date, region, and advertiser dimensions.
You are asked to redesign the storage layout and downstream ETL strategy, and clearly explain the difference between partitioning and bucketing in Hive or Spark. Your design should show when each technique is appropriate, how they interact with query patterns, and how to avoid small-file and skew issues.
event_date, region, and joins on advertiser_id.