Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started
Dataford
Popular roles
Software EngineerData AnalystData ScientistData EngineerBusiness AnalystAI EngineerMachine Learning EngineerProduct Manager
Browse
Browse All RolesEvery role hub, from analyst to MLBrowse All CompaniesCompany-specific interview loopsAll Interview GuidesThe full guide library
Top questions by role
Software EngineerData AnalystData ScientistData EngineerBusiness AnalystAI EngineerMachine Learning EngineerProduct Manager
Top questions by skill
SQLPythonStatisticsMachine LearningA/B TestingSystem DesignGenerative AIProduct SenseMetricsBehavioral
Browse all questions →Try a mock interview
Experiences
Practice
Mock InterviewsTimed interview simulations with feedbackSuccess PathYour 6-week structured planModulesCurated lessons by topicWebinarsTalks from ex-Big Tech data leadsPlaygroundA free-form scratch editor
Learn
BlogInterview strategy and career adviceTech Job Market ReportHiring trends across data and AI rolesFor UniversitiesDataford for career centersAbout DatafordWho we are and how we build
Pricing
Build my plan

Clinical Document QA: Fine-Tune vs RAG

MediumNLP00:00
Practice interviewer
In session
5 left
00:00

Your question is Clinical Document QA: Fine-Tune vs RAG. Take a moment with it on the right.

Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).

You need to log in / sign up to chat or submit.

Problem

Business Context

Northstar Health wants a production system that answers clinician questions over internal policies, care pathways, discharge instructions, and trial protocols. You must design and evaluate whether a fine-tuned LLM, a retrieval-augmented generation (RAG) pipeline, or a hybrid approach is the better fit for internal clinical document querying.

Data

  • Corpus: 420,000 internal clinical documents across PDF, DOCX, HTML, and scanned OCR text
  • Text length: 1 paragraph to 180 pages; median chunkable length is 320 words
  • Language: English only, but includes abbreviations, ICD/CPT codes, drug names, and templated sections
  • Question set: 18,000 historical clinician queries with 6,500 answerable QA pairs and document citations
  • Label distribution: ~55% fact lookup, 25% policy/procedure questions, 12% multi-document synthesis, 8% unanswerable or outdated

Success Criteria

The system is good enough if it achieves grounded answer quality with citation support, answer faithfulness above 90% on answerable questions, top-5 retrieval recall above 92% for RAG-style systems, and median end-to-end latency below 2 seconds for interactive use.

Constraints

  • HIPAA-compliant deployment in a private VPC; no external API calls
  • Documents update daily, with versioned policies and retired content
  • Answers must cite source passages and avoid unsupported medical claims
  • Budget supports one 24GB GPU for training/inference plus a managed vector store

Requirements

  1. Propose a fine-tuned LLM approach and a RAG pipeline for this corpus.
  2. Compare trade-offs in freshness, hallucination risk, maintenance cost, latency, and citation quality.
  3. Build a realistic preprocessing pipeline for ingestion, chunking, metadata extraction, and de-identification.
  4. Implement a baseline modern Python solution for document indexing, retrieval, generation, and evaluation.
  5. Define an offline evaluation plan and recommend which architecture you would ship first, with justification.