Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started

Debug Memory Leaks and Bottlenecks

MediumCoding00:00
Practice interviewer
In session
5 left
00:00

Your question is Debug Memory Leaks and Bottlenecks. Take a moment with it on the right.

Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).

You need to log in / sign up to chat or submit.

Problem

Nexla runs a no-code data integration platform: every connector sync pulls records from a customer's source system, checks them against a schema cache, and writes an audit trail. A support ticket came in about a connector pod that starts fine but has its memory climb steadily over a few hours until it OOMs and gets restarted by the orchestrator, and syncs on large batches are also noticeably slower than they used to be.

Here is the function on-call pulled from the sync worker:

import time

_schema_cache = {}


def sync_connector_batch(connector_id, records, audit_path="audit.log"):
    seen_ids = []
    summary_lines = ""

    for record in records:
        request_id = f"{connector_id}-{record['id']}-{time.time_ns()}"
        raw_response = fetch_schema(connector_id, record)
        _schema_cache[request_id] = raw_response

        if record["id"] in seen_ids:
            continue
        seen_ids.append(record["id"])

        audit_file = open(audit_path, "a")
        audit_file.write(f"synced {record['id']}
")

        summary_lines += f"{record['id']}: ok
"

    return summary_lines


def fetch_schema(connector_id, record):
    # pretend this calls the connector's remote API
    return {"connector_id": connector_id, "record": record, "fetched_at": time.time()}

Explain in your own words why memory keeps growing and why large batches get slower over time, and what you would change. You are not editing the file, just walking an interviewer through the diagnosis.