Your question is Adversarial Test for Prompt Safety. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
How would you design an adversarial test set to identify and mitigate prompt circumvention in an LLM?
Describe a practical dataset, labeling scheme, attack taxonomy, split strategy, and evaluation harness. Explain how you would prevent template leakage, measure robustness on unseen attacks, select operating thresholds, and connect findings to mitigation such as NVIDIA NeMo Guardrails or supervised fine-tuning. Include production monitoring and an experiment that distinguishes genuine safety improvement from attackers adapting to the test set.