Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started
Optimizing Window Functions at Scale
00:00
5 left

Optimizing Window Functions at Scale

HardSQL · PostgreSQL

Problem

How do you optimize a query that utilizes a window function on a massive, distributed dataset at Plaid?

Use the provided PostgreSQL tables and write the query for completed January transactions. The solution should reduce the data processed by the window calculation while producing deterministic results.

Output

  1. One row per account and transaction date with completed transactions
  2. Columns: account_label, transaction_date, daily_net_amount, running_balance, and prior_running_balance
  3. Order by account_label, then transaction_date

Schema

plaid_accounts
ColumnTypeDescription
account_idPKINTUnique Plaid account identifier
account_labelVARCHAR(100)Display label for the account
regionVARCHAR(30)Account region
plaid_transactions
ColumnTypeDescription
transaction_idPKINTUnique transaction identifier
account_idINTReferenced Plaid account
transaction_dateDATEPosting date of the transaction
amountNUMERIC(12,2)Signed transaction amount
statusVARCHAR(20)Transaction processing status
Tablesplaid_accountsplaid_transactions
Interviewer

Your question is Optimizing Window Functions at Scale. Start with the requirements and the two tables in the Question tab.

Run and submit as often as you like. When you're ready, talk me through your approach or go straight to the code.

You need to log in / sign up to run or submit.
CodePostgreSQL
You need to log in / sign up to run or submit.Ln 1
Run your query to see results here.