Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started
Dataford
Popular roles
Software EngineerData AnalystData ScientistData EngineerBusiness AnalystAI EngineerMachine Learning EngineerProduct Manager
Browse
Browse All RolesEvery role hub, from analyst to MLBrowse All CompaniesCompany-specific interview loopsAll Interview GuidesThe full guide library
Top questions by role
Software EngineerData AnalystData ScientistData EngineerBusiness AnalystAI EngineerMachine Learning EngineerProduct Manager
Top questions by skill
SQLPythonStatisticsMachine LearningA/B TestingSystem DesignGenerative AIProduct SenseMetricsBehavioral
Browse all questions →Try a mock interview
Experiences
Practice
Mock InterviewsTimed interview simulations with feedbackSuccess PathYour 6-week structured planModulesCurated lessons by topicWebinarsTalks from ex-Big Tech data leadsPlaygroundA free-form scratch editor
Learn
BlogInterview strategy and career adviceTech Job Market ReportHiring trends across data and AI rolesFor UniversitiesDataford for career centersAbout DatafordWho we are and how we build
Pricing
Build my plan
SQL-like Data Functions On Demand
00:00
5 left

SQL-like Data Functions On Demand

HardSQL · PostgreSQL

Problem

Given a database schema and documentation, implement functions on the spot using the documentation.

Asked in the learning stage. VO2; repeated in the HM technical portion.

Implement the documented PostgreSQL function get_pipeline_health(workspace_id, month_start). Return one row for every active pipeline in the workspace. Include only runs from month_start through the final day of that month. A pipeline with no runs must still appear. failure_rate_pct is failed runs divided by all runs, rounded to two decimal places. Health is no runs for zero runs, healthy for a failure rate at most 10%, and attention otherwise.

Output

  1. Columns: pipeline_id, pipeline_name, total_runs, successful_runs, failure_rate_pct, health_status
  2. One row per active pipeline in workspace 1 for February 2025
  3. Order by pipeline_name ascending

Use the documented function contract and schema below.

Schema

workspaces
ColumnTypeDescription
workspace_idPKINTUnique workspace identifier
workspace_nameVARCHAR(100)Workspace display name
pipelines
ColumnTypeDescription
pipeline_idPKINTUnique pipeline identifier
workspace_idINTWorkspace containing the pipeline
pipeline_nameVARCHAR(120)Pipeline display name
owner_teamVARCHAR(100)Team responsible for the pipeline
is_activeBOOLEANWhether the pipeline should be included
pipeline_runs
ColumnTypeDescription
run_idPKINTUnique pipeline run identifier
pipeline_idINTPipeline executed by the run
run_started_atTIMESTAMPRun start timestamp
statusVARCHAR(20)Run outcome, normally success or failure
Tablesworkspacespipelinespipeline_runs
Interviewer

Your question is SQL-like Data Functions On Demand. Start with the requirements and the three tables in the Question tab.

Run and submit as often as you like. When you're ready, talk me through your approach or go straight to the code.

You need to log in / sign up to run or submit.
CodePostgreSQL
Sign up free to run your codeLog inLn 1
Run your query to see results here.