Your question is Minimize Harmful LLM Outputs. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
How would you design a language model that minimizes harmful outputs while still being useful and expressive? Consider the complete system, including data curation, post-training, inference-time safeguards, refusal behavior, and evaluation. Explain how you would handle ambiguity, adversarial prompts, distribution shift, cultural variation, and the tradeoff between safety and over-refusal.