Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started
Dataford
Popular roles
Software EngineerData AnalystData ScientistData EngineerBusiness AnalystAI EngineerMachine Learning EngineerProduct Manager
Browse
Browse All RolesEvery role hub, from analyst to MLBrowse All CompaniesCompany-specific interview loopsAll Interview GuidesThe full guide library
Top questions by role
Software EngineerData AnalystData ScientistData EngineerBusiness AnalystAI EngineerMachine Learning EngineerProduct Manager
Top questions by skill
SQLPythonStatisticsMachine LearningA/B TestingSystem DesignGenerative AIProduct SenseMetricsBehavioral
Browse all questions →Try a mock interview
Experiences
Practice
Mock InterviewsTimed interview simulations with feedbackSuccess PathYour 6-week structured planModulesCurated lessons by topicWebinarsTalks from ex-Big Tech data leadsPlaygroundA free-form scratch editor
Learn
BlogInterview strategy and career adviceTech Job Market ReportHiring trends across data and AI rolesFor UniversitiesDataford for career centersAbout DatafordWho we are and how we build
Pricing
Build my plan
SQL, Python, and Pandas/Spark Case
00:00
5 left

SQL, Python, and Pandas/Spark Case

MediumSQL · PostgreSQL

Problem

How would you solve a case study that combines SQL, Python, and Pandas or Spark?

Use the supplied relational data to produce the SQL result, then describe how you would reproduce and validate the same result in Python with Pandas or Spark. Include assumptions, data-quality checks, and how you would handle missing values.

Output

  1. One row per customer with customer_id, customer_name, completed_order_count, total_completed_amount, and latest_completed_order_date
  2. Include customers without completed orders, using zero totals where appropriate
  3. Order by total_completed_amount descending, then customer_id ascending

Schema

customers
ColumnTypeDescription
customer_idPKINTUnique customer identifier
customer_nameVARCHAR(100)Customer display name
orders
ColumnTypeDescription
order_idPKINTUnique order identifier
customer_idINTCustomer associated with the order
statusVARCHAR(20)Order lifecycle status
amountNUMERIC(10,2)Order amount
order_dateDATEDate the order was placed
Tablescustomersorders
Interviewer

Your question is SQL, Python, and Pandas/Spark Case. Start with the requirements and the two tables in the Question tab.

Run and submit as often as you like. When you're ready, talk me through your approach or go straight to the code.

You need to log in / sign up to run or submit.
CodePostgreSQL
Sign up free to run your codeLog inLn 1
Run your query to see results here.