Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started

Evaluate Safe LLM Response Quality

Easy
Model EvaluationPrecisionAccuracyRecall

Problem

Context

HealthAssist AI is a customer-facing LLM that answers general wellness questions and drafts support responses for a telehealth platform. The team added a safety layer to block harmful or non-compliant outputs, but users now report that some safe questions are being refused while a smaller set of risky answers still slip through.

Current Performance

MetricValidation SetTargetNotes
Safe response precision0.96>= 0.95Most allowed answers are actually safe
Safe response recall0.78>= 0.90Many safe queries are unnecessarily blocked
Harmful output rate1.8%< 0.5%Too many unsafe responses still reach users
Refusal rate24%10-15%Over-refusal hurts usability
F1 score (safe vs unsafe)0.86>= 0.92Overall balance is weak
Calibration error0.11< 0.05Risk scores are poorly aligned to actual risk

The Problem

The model appears conservative on benign prompts but still misses some genuinely unsafe outputs. Leadership wants a practical evaluation plan that improves both safety and answer quality without making the assistant unusable.

Requirements

  1. Interpret what the current metrics imply about safety vs usability tradeoffs.
  2. Identify likely failure modes causing both over-refusal and unsafe leakage.
  3. Propose an evaluation framework covering offline tests, human review, and launch gates.
  4. Recommend threshold, calibration, and policy improvements.
  5. Define how you would monitor post-launch drift and regressions.

Constraints

  • Harmful medical advice is high-severity and must be minimized.
  • Excessive refusals reduce user trust and support deflection.
  • Human review budget covers only 2,000 prompts per week.
  • Any production change must be explainable to compliance and policy teams.

You are practicing as a guest. Sign up free to get your answer graded with AI feedback. Your draft stays right here.

Sign up freeI have an account
Sign up to unlock solutions
Next questions
Evaluate Safe Helpful AI ResponsesHardPersistent SystemsEvaluate GenAI Quality and SafetyEasySpotOn: CorporateEvaluate a Grounded Support AssistantMedium