Skip to main content

A request to a model is not one block of text. It carries a system prompt, the user’s message, earlier turns, tool definitions, tool arguments, and tool results — and each of those has a different author and a different risk. TrustGuard splits the payload by role and sends each part only to the detectors built for it.

This page is the map. Use it to predict which detector will fire on which part of a turn, and to understand why a rule you expected to match did not.

At a glance

Input is the request sent to the model or MCP server. Output is what comes back. The phase is the one you pick on the policy rule. Two consequences are worth knowing before you write rules:
  • A tool result on an LLM call is Input, not Output. The model asks for a tool in one response; your agent runs it and sends the result back in the next request. That is where TrustGuard sees it. On MCP, the result is the server’s response, so it arrives on Output.
  • No single detector covers everything. Prompt Guard never reads tool results, and Indirect Prompt Injection never reads what the user typed. An agent that calls tools needs both. See Agent & MCP security.

Detector by detector

Prompt Guard

Reads text that is meant to steer the model:
  • System and developer instructions, treated as prompt-side content.
  • The user’s message, including the text blocks of a multimodal message.
  • Tool definitions on an LLM request — each tool’s description, and the descriptions of its parameters. An instruction hidden in a description is how a poisoned tool attacks before it is ever called.
  • Assistant text, on Output and when it is replayed in a later request’s history.
Prompt Guard uses a different model for each side. Prompt-side text is scored for jailbreak attempts by whoever wrote it; assistant text is scored for a model that has already been jailbroken. You configure one detector; TrustGuard picks the model by where the text came from. It does not read tool results, tool-call arguments, or reasoning fields.

Toxicity Detection and Custom Moderation

Both read conversational text in either direction:
  • The user’s message on Input.
  • Assistant text on Output, and when it is replayed in history.
They do not read the system prompt, tool definitions, or tool traffic.

Data Loss Prevention

Reads text where personal data and secrets actually leave or enter:
  • The user’s message on Input.
  • Tool-call argument values on Input — the values your agent is about to send to a tool, such as an email address passed to a CRM lookup.
  • Assistant text on Output.
It is the only detector that can Transform, masking a matched value in place.

Indirect Prompt Injection

Reads content that came from a tool rather than from anyone in the conversation:
  • Tool results returned to the model. Structured results are walked recursively and every string is scored, error messages included.
  • MCP tools/list results — tool descriptions, plus the titles, descriptions, and defaults inside each input schema.
  • MCP tool results on Output.
It does not read the user’s message, the assistant’s replies, or the system prompt. A turn with no tool involved produces no Indirect Prompt Injection findings.

URL Analyzer and Document Analyzer

These two do not score the user’s message directly. They pull external content out of it and score that:
  • URL Analyzer finds each URL in the user’s message, fetches the page with SSRF protections, and scores the returned text.
  • Document Analyzer extracts the text of each attached file — running OCR on scans and images — and scores that.
The extracted content goes to the two detectors you linked when you created the analyzer: an injection check and a PII check. The original user text is still scored by Prompt Guard as usual. A URL is not its content, so the two are evaluated separately.

What is never scanned

Provider field reference

When you send a full provider body to POST /v1/evaluate — or when a gateway sends it for you — TrustGuard extracts the parts above from these fields. Tool-call arguments live in tool_calls[].function.arguments (OpenAI Chat, Cohere), input[type=function_call].arguments (OpenAI Responses), tool_use.input (Anthropic), and functionCall.args (Gemini). Those are the values Data Loss Prevention reads.

Example: one tool-using turn

An OpenAI Chat Completions agent is asked to summarise a web page. The first request carries a system prompt, the user’s message, and one tool:
The model answers with a tool call and no text, so Output has nothing to score. The agent runs the tool and sends the result back in the next request:
The final response carries the summary in choices[0].message.content, which Prompt Guard, Toxicity Detection, Custom Moderation, and Data Loss Prevention read on Output.

MCP traffic

On MCP, TrustGuard reads the JSON-RPC messages between the agent and the server. Scoring tools/list is what catches tool poisoning: a server that advertises a tool with instructions in its description is flagged before any tool is called. Only collectors that see the tool listing can do this — see How it works for which ones.