A request to a model is not one block of text. It carries a system prompt, the user’s message, earlier turns, tool definitions, tool arguments, and tool results — and each of those has a different author and a different risk. TrustGuard splits the payload by role and sends each part only to the detectors built for it.
This page is the map. Use it to predict which detector will fire on which part of a turn, and to understand why a rule you expected to match did not.At a glance
Input is the request sent to the model or MCP server. Output is what comes back. The phase is the one you pick on the policy rule.
Two consequences are worth knowing before you write rules:
- A tool result on an LLM call is Input, not Output. The model asks for a tool in one response; your agent runs it and sends the result back in the next request. That is where TrustGuard sees it. On MCP, the result is the server’s response, so it arrives on Output.
- No single detector covers everything. Prompt Guard never reads tool results, and Indirect Prompt Injection never reads what the user typed. An agent that calls tools needs both. See Agent & MCP security.
Detector by detector
Prompt Guard
Reads text that is meant to steer the model:- System and developer instructions, treated as prompt-side content.
- The user’s message, including the text blocks of a multimodal message.
- Tool definitions on an LLM request — each tool’s description, and the descriptions of its parameters. An instruction hidden in a description is how a poisoned tool attacks before it is ever called.
- Assistant text, on Output and when it is replayed in a later request’s history.
Toxicity Detection and Custom Moderation
Both read conversational text in either direction:- The user’s message on Input.
- Assistant text on Output, and when it is replayed in history.
Data Loss Prevention
Reads text where personal data and secrets actually leave or enter:- The user’s message on Input.
- Tool-call argument values on Input — the values your agent is about to send to a tool, such as an email address passed to a CRM lookup.
- Assistant text on Output.
Indirect Prompt Injection
Reads content that came from a tool rather than from anyone in the conversation:- Tool results returned to the model. Structured results are walked recursively and every string is scored, error messages included.
- MCP
tools/listresults — tool descriptions, plus the titles, descriptions, and defaults inside each input schema. - MCP tool results on Output.
URL Analyzer and Document Analyzer
These two do not score the user’s message directly. They pull external content out of it and score that:- URL Analyzer finds each URL in the user’s message, fetches the page with SSRF protections, and scores the returned text.
- Document Analyzer extracts the text of each attached file — running OCR on scans and images — and scores that.
What is never scanned
Provider field reference
When you send a full provider body toPOST /v1/evaluate
— or when a gateway sends it for you — TrustGuard extracts the parts above from
these fields.
Tool-call arguments live in
tool_calls[].function.arguments (OpenAI Chat,
Cohere), input[type=function_call].arguments (OpenAI Responses), tool_use.input
(Anthropic), and functionCall.args (Gemini). Those are the values Data Loss
Prevention reads.
Example: one tool-using turn
An OpenAI Chat Completions agent is asked to summarise a web page. The first request carries a system prompt, the user’s message, and one tool:
The model answers with a tool call and no text, so Output has nothing to score.
The agent runs the tool and sends the result back in the next request:
The final response carries the summary in
choices[0].message.content, which
Prompt Guard, Toxicity Detection, Custom Moderation, and Data Loss Prevention read
on Output.
MCP traffic
On MCP, TrustGuard reads the JSON-RPC messages between the agent and the server.
Scoring
tools/list is what catches tool poisoning: a server that advertises a
tool with instructions in its description is flagged before any tool is called.
Only collectors that see the tool listing can do this — see
How it works for which ones.