Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started

Evaluate GenAI Quality and Safety

Easy
Model EvaluationPrecisionAccuracyRecallAsked 1 times

Problem

Context

BrightAssist is deploying a customer-support generative AI model that answers billing and account questions in a fintech app. In offline evaluation, the model produces fluent responses, but the trust and safety team found cases of incorrect financial advice, policy violations, and unsafe escalation handling.

Current Performance

MetricCurrent ModelTargetNotes
Helpfulness pass rate78%85%Human-rated on 1,200 prompts
Factual accuracy74%90%Grounded against internal KB
Policy safety pass rate96.2%99.0%Includes harmful/regulated content checks
Hallucination rate11.5%5.0%Unsupported claims in final answer
Escalation recall68%90%Cases that should be handed to a human
Over-refusal rate14%7%Safe requests incorrectly declined
Avg. response latency2.1s2.5sWithin SLA

The Problem

Leadership wants a practical evaluation framework that measures both answer quality and safety before launch. The current metrics suggest the model is usable for simple requests but unreliable in high-risk situations, especially where escalation or factual grounding is required.

Requirements

  1. Define an evaluation approach covering both quality and safety.
  2. Interpret the current metrics and identify the biggest launch risks.
  3. Recommend how to segment evaluation by use case, risk level, and prompt type.
  4. Propose thresholding or gating rules for launch readiness.
  5. Suggest concrete improvements to raise quality without weakening safety.

Constraints

  • The product handles regulated financial topics.
  • Human review capacity is limited to 8% of daily conversations.
  • Launch cannot increase average latency above 2.5 seconds.
  • A severe unsafe response is considered more costly than an unnecessary refusal.
Practicing as: Data Scientist interview at Persistent Systems

Hi, I'll play your Persistent Systems interviewer for the Data Scientist role. Answer the question above like we're in the room, and I'll respond the way a real interviewer would.

You are practicing as a guest. Sign up free to get your answer graded with AI feedback. Your draft stays right here.

Sign up freeI have an account
Sign up to unlock solutions
Persistent Systems Data Scientist Interview Questions
Next questions
Evaluate Safe Helpful AI ResponsesHardEvaluate Safe LLM Response QualityEasyOpenAIAlignment vs Evaluation TradeoffEasy