Purpose
Off-topic testing is essential for:- Ensuring models stay within their intended scope
- Preventing engagement with sensitive or unauthorized topics
- Maintaining brand consistency
- Avoiding liability from inappropriate responses
Content Categories
Scope Boundaries
Sensitive Topics
Dataset vs Prompt probes
The catalog builder instantiates*DatasetProbe classes and evaluates with TrueFalseEvaluator. Each subcategory also has a *PromptProbe (for example CompetitorsPromptProbe).
Dataset probes accept sampling, use_jailbreaks, max_jailbreak_per_objective, and translate_into_language. The builder maps language → translate_into_language and num_test_cases → max_objectives_per_category. See StaticDatasetProbe.
Catalog scenarios attach EU AI Act Art. 15, OWASP AITG-APP-05 / LLM05, MITRE ATLAS AML.T0054, and ISO/IEC 42001 tags.
How It Works
The evaluator uses a binary scoring system:- Pass: The model appropriately stays within scope or redirects
- Fail: The model engages with off-topic content it should avoid
Usage Example
When to Use
Use off-topic testing when you need to:- Define and enforce content boundaries
- Ensure models stay on-brand
- Prevent engagement with sensitive topics
- Test domain-specific assistants
- Validate content policy compliance