Skip to main content
Detectors find things. A policy is what decides whether anything happens about it — and it is the only place in TrustGuard where that decision is made. A policy holds three things: gates that filter traffic before anything is inspected, detector rules that say which detector runs and what its finding costs, and one switch between Observe and Enforce.

The six tabs

Opening a policy opens a side panel: In the policy list, the Gates, Detectors, and Collectors columns count what is on those tabs, and Mode shows Observe or Enforce.

Enforcement mode

The switch is on Basics, and also in the panel header so you can flip it without digging. Run every new policy in Observe first. Check Activity, use the Test tab, and only then switch. This is not ceremony — a policy that looks obviously correct can still be noisy against real traffic, and Observe is where you learn that without anyone noticing.

Gates

Gates match on request attributes — who is calling, which model, which tool — never on the text of the prompt. They run before any detector, which makes them the cheap filter: traffic a gate skips costs you nothing. Each gate is a Rule name, a set of Conditions joined by And or Or, and a Then action. All gate rules evaluate, and the most restrictive action wins — they are not a first-match chain, so adding a rule can only ever tighten a policy, never loosen one. Ask depends on the integration having somewhere to ask. Claude Code and Cursor show a permission prompt for tool calls. GitHub Copilot prompts on interactive tool calls and treats Ask as denied in cloud jobs. Codex has no dialog at all: it allows the call and adds the approval message to the agent’s context. Full mapping in How it works.

What you can match on

The Attribute picker groups everything you can gate on: That is the whole list — the picker is a closed menu, so anything not in it cannot be gated on. The picker shows a one-line description of each as you hover it, so this table is a map rather than something to memorise. A few worth calling out:
  • Text analyzed (bytes) is how much text the detectors actually read, which is what their cost scales with — the attribute to gate on if you are managing spend.
  • Tool results already cleared counts results skipped because they were checked and allowed earlier, so a rising number means your policy is doing redundant work.
  • Agent published separates a live agent from one still being built, which is usually the difference between enforcing and observing.
  • Tool arguments matches the whole argument object as a single string. There is no way to gate on one named argument — if you need that, match on Tool name and let a detector inspect the content.
Operators: Equals, Not equals, Greater than, Less than, Contains, Does not contain, In, Matches. In takes a list of values; Matches takes a regular expression.

Source application values

Source application tells you which surface a request came from, and the values are fixed:
claude-code and claude-code-plugin are different surfaces. A gate on claude-code does not match the laptop plugin, which is the single most common reason a gate silently never fires.

Detectors

The Detectors tab is split by Evaluation phase:
  • Input — traffic going to the model or tool.
  • Output — traffic coming back.
Only the phase matching the traffic runs. That sounds obvious but has a practical edge: a collector that only ever sees prompts has nothing to evaluate on Output, so Output rules there will never fire no matter how they are written. Which phases a collector can see is on its integration page. Each rule picks a detector, an Action, and optional Conditions — the same attributes as gates. As with gates, every rule evaluates and the most restrictive action wins.

When several things fire

One request can trip a gate and three detectors. TrustGuard reduces all of it to a single outcome, worst first:
Allow means nothing fired, or everything that fired was waived. In Observe, the outcome never goes past Monitor — which is exactly what makes Observe safe.

Test

Test runs against the saved policy. Unsaved gates and detector rules are not applied, so save before you test or you will be testing the old version.
  1. Pick a Test direction — Input or Output. It has to match the phase of the rule you are trying to exercise.
  2. Paste a sample input or output, pick one of the built-in presets — Prompt Injection, PII, Jailbreak, Secrets, Toxicity — or upload a file.
  3. Fill in Extra parameters, the same attributes gates use. They are prefilled from the policy’s own gates, and Skip, Block and Ask rules only match when the attribute they gate on is actually set here — so a rule keyed on Tool name will look broken until you fill it in.
  4. Run test and read the outcome and its findings.
Test does not write to Activity and does not count against your plan’s rate limit, so use it freely.

Collectors

A policy does nothing until traffic reaches it. On the Collectors tab, attach collectors and pick a Routing mode:
  • Default — the fallback for all of that collector’s traffic.
  • Consumer ID — only requests from that consumer, which is how one collector sends different consumers to different policies.
You can bind the same routing from the collector’s own Policies tab.
A collector with no matching policy leaves that traffic unguarded. TrustGuard returns allow and inspects nothing — it does not warn you. This is the last step, and the one people forget.

History

History records every create, update, and delete on the policy, its gates, its detector rules, and its collector routing. When enforcement changes and nobody remembers why, this is where the answer is.