The six tabs
Opening a policy opens a side panel:
In the policy list, the Gates, Detectors, and Collectors columns count
what is on those tabs, and Mode shows Observe or Enforce.
Enforcement mode
The switch is on Basics, and also in the panel header so you can flip it without digging.
Run every new policy in Observe first. Check Activity, use the Test
tab, and only then switch. This is not ceremony — a policy that looks obviously
correct can still be noisy against real traffic, and Observe is where you learn
that without anyone noticing.
Gates
Gates match on request attributes — who is calling, which model, which tool — never on the text of the prompt. They run before any detector, which makes them the cheap filter: traffic a gate skips costs you nothing. Each gate is a Rule name, a set of Conditions joined by And or Or, and a Then action.
All gate rules evaluate, and the most restrictive action wins — they are not
a first-match chain, so adding a rule can only ever tighten a policy, never
loosen one.
Ask depends on the integration having somewhere to ask. Claude
Code and Cursor show a
permission prompt for tool calls. GitHub
Copilot prompts on interactive tool calls and
treats Ask as denied in cloud jobs. Codex has no dialog at
all: it allows the call and adds the approval message to the agent’s context.
Full mapping in How it works.
What you can match on
The Attribute picker groups everything you can gate on:
That is the whole list — the picker is a closed menu, so anything not in it
cannot be gated on.
The picker shows a one-line description of each as you hover it, so this table is
a map rather than something to memorise. A few worth calling out:
- Text analyzed (bytes) is how much text the detectors actually read, which is what their cost scales with — the attribute to gate on if you are managing spend.
- Tool results already cleared counts results skipped because they were checked and allowed earlier, so a rising number means your policy is doing redundant work.
- Agent published separates a live agent from one still being built, which is usually the difference between enforcing and observing.
- Tool arguments matches the whole argument object as a single string. There is no way to gate on one named argument — if you need that, match on Tool name and let a detector inspect the content.
Source application values
Source application tells you which surface a request came from, and the values are fixed:Detectors
The Detectors tab is split by Evaluation phase:- Input — traffic going to the model or tool.
- Output — traffic coming back.
As with gates, every rule evaluates and the most restrictive action wins.
When several things fire
One request can trip a gate and three detectors. TrustGuard reduces all of it to a single outcome, worst first:Test
Test runs against the saved policy. Unsaved gates and detector rules are not applied, so save before you test or you will be testing the old version.- Pick a Test direction — Input or Output. It has to match the phase of the rule you are trying to exercise.
- Paste a sample input or output, pick one of the built-in presets — Prompt Injection, PII, Jailbreak, Secrets, Toxicity — or upload a file.
- Fill in Extra parameters, the same attributes gates use. They are prefilled from the policy’s own gates, and Skip, Block and Ask rules only match when the attribute they gate on is actually set here — so a rule keyed on Tool name will look broken until you fill it in.
- Run test and read the outcome and its findings.
Collectors
A policy does nothing until traffic reaches it. On the Collectors tab, attach collectors and pick a Routing mode:- Default — the fallback for all of that collector’s traffic.
- Consumer ID — only requests from that consumer, which is how one collector sends different consumers to different policies.