Each label set gives a request at most one of its labels, or none — unlabeled for
that set — when no label clearly applies. Sets are independent: a request can be
negative in Sentiment, complaint in Intent and billing in Domain.
Labels are observability only. They are produced after the request, in the background, and
never reach policies, routing or the response.
Set it up
1
Turn it on for the gateway
In Settings → Agent Gateway → Traffic labels, switch on Label traffic and pick
the Registry and Model that run the classification. The registry must be an
LLM registry of the same gateway that holds its own credentials: classification runs in
the background with no client key to forward, so pass-through and OAuth2 registries are
not offered.
2
Create label sets
Under Label sets on the same tab, add each set: a name, its instructions, and its
labels — at least two — each with a name and a description. Descriptions are what the
classifier reads to tell labels apart, so a line on each pays off.
3
Assign them to applications
On the application’s Label sets tab, pick the sets that apply to its traffic. Only
an application with a model provider (an LLM endpoint) is labeled: tool and agent
traffic is not.
Limits
When an application shows Out of sync
Label sets are kept in the console and pushed to the gateway whenever you assign them, edit a set or change an application’s endpoints. If that push fails, the change is saved but the gateway keeps labeling with the previous sets, and the application’s Label sets tab shows Out of sync with a Retry. Editing or deleting a set that several applications use pushes it to each of them; the settings tab says how many did not sync.What gets classified
Only chat requests are offered — Chat Completions, Responses, Anthropic Messages, Gemini, Cohere chat. Embeddings, rerank, files, images and audio are not. Requests a guardrail blocks are labeled too: classification sees the request as the client sent it. The classifier reads the latest user messages of the conversation, up to the messages window, and at most the last 10,000 characters of them. System prompts and assistant replies are never sent. Where those messages come from depends on the API:
For the last row the gateway needs to know which conversation a turn belongs to; see
Grouping a conversation.
That record is kept encrypted for an hour after the last turn; after a longer pause the
next turn is classified on its own messages.
All the application’s label sets are classified in one call to your model per request.
Identical text against the same sets, registry and model is answered from a cache instead of
calling the model again.
Cost and privacy
Classification is a normal completion billed by your provider on the registry you picked: one call per classified request, its input being the label sets plus the window. Use Sampling rate to classify a fraction of the traffic on busy applications, and a small, fast model — the task needs no reasoning depth. What leaves the gateway, and where it goes:- To your classifier registry: the window’s user messages and the application’s label sets.
- To the gateway’s own queue: the same, until classified, then deleted.
- To analytics: per request and label set, the label (or none), plus the model, the registry, token usage and latency. Never the prompt.
Read the results
Analytics → Labels
Pick a Label set in the toolbar, next to the application filter — or All labels, the default, to see every set at once.- Label Distribution — requests per label over time, with an optional Unlabeled series.
- Labels — each label’s requests and share, ending with the Unlabeled row.
- Users — sessions, unique users, new users and sessions per user among the classified requests. A user is the end user a client declared, otherwise the authenticated principal.
Filtering by an application narrows every block to that application’s traffic.
A request’s detail
Opening a request in Activity shows one line per label set it was classified against —Sentiment · negative, or — when the set gave none — and the model that classified it.
Classification finishes a few seconds after the request, so a very recent request may not
show it yet.
Troubleshooting
Related
- Applications: where label sets are assigned
- Settings: the gateway’s Traffic labels tab
- End-user attribution: who a user is, and how conversations are grouped