How would you solve a case study that combines SQL, Python, and Pandas or Spark?
Use the supplied relational data to produce the SQL result, then describe how you would reproduce and validate the same result in Python with Pandas or Spark. Include assumptions, data-quality checks, and how you would handle missing values.
customer_id, customer_name, completed_order_count, total_completed_amount, and latest_completed_order_datetotal_completed_amount descending, then customer_id ascending| Column | Type | Description |
|---|---|---|
| customer_idPK | INT | Unique customer identifier |
| customer_name | VARCHAR(100) | Customer display name |
| Column | Type | Description |
|---|---|---|
| order_idPK | INT | Unique order identifier |
| customer_id | INT | Customer associated with the order |
| status | VARCHAR(20) | Order lifecycle status |
| amount | NUMERIC(10,2) | Order amount |
| order_date | DATE | Date the order was placed |