Your question is Evaluate LLM Safety Metrics. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are evaluating a conversational model used in a chat product. The team wants a clear way to measure whether responses stay safe, fair, and aligned with user intent across different prompt types and user groups.
How would you develop metrics to assess model toxicity, bias, and alignment?