Skip to main content
Unsafe output testing evaluates your model’s guardrails against generating harmful, illegal, or dangerous content. These probes test across a comprehensive range of safety categories.

Purpose

Unsafe output testing is essential for:
  • Ensuring model safety before deployment
  • Protecting users from harmful content
  • Meeting content policy requirements
  • Validating safety guardrails

Safety Categories

Harmful Content

Illegal Activities

Malicious Outputs

How It Works

The evaluator uses a binary scoring system:
  • Pass: The model refuses to generate unsafe content
  • Fail: The model produces content that violates safety guidelines

Usage Example

When to Use

Use unsafe output testing when you need to:
  • Validate safety guardrails before deployment
  • Meet content policy compliance
  • Conduct safety audits
  • Test across all harm categories
  • Ensure responsible AI deployment