Skip to main content
Request counts and token costs say how much an application is used, not what it is used for. Traffic labels answer that. You describe the categories you care about as label sets, assign them to applications, and an LLM you choose classifies each chat request against them. The results land in Analytics → Labels and in each request’s detail. A label set is a name, optional instructions saying what it classifies, and the labels to choose from, each with an optional description: Each label set gives a request at most one of its labels, or none — unlabeled for that set — when no label clearly applies. Sets are independent: a request can be negative in Sentiment, complaint in Intent and billing in Domain.
Labels are observability only. They are produced after the request, in the background, and never reach policies, routing or the response.

Set it up

1

Turn it on for the gateway

In Settings → Agent Gateway → Traffic labels, switch on Label traffic and pick the Registry and Model that run the classification. The registry must be an LLM registry of the same gateway that holds its own credentials: classification runs in the background with no client key to forward, so pass-through and OAuth2 registries are not offered.
2

Create label sets

Under Label sets on the same tab, add each set: a name, its instructions, and its labels — at least two — each with a name and a description. Descriptions are what the classifier reads to tell labels apart, so a line on each pays off.
3

Assign them to applications

On the application’s Label sets tab, pick the sets that apply to its traffic. Only an application with a model provider (an LLM endpoint) is labeled: tool and agent traffic is not.
A request is classified only when both are true: labeling is on for its gateway, and its application has at least one label set. Turning Label traffic off keeps the configuration and the label sets; it just stops classifying.

Limits

When an application shows Out of sync

Label sets are kept in the console and pushed to the gateway whenever you assign them, edit a set or change an application’s endpoints. If that push fails, the change is saved but the gateway keeps labeling with the previous sets, and the application’s Label sets tab shows Out of sync with a Retry. Editing or deleting a set that several applications use pushes it to each of them; the settings tab says how many did not sync.

What gets classified

Only chat requests are offered — Chat Completions, Responses, Anthropic Messages, Gemini, Cohere chat. Embeddings, rerank, files, images and audio are not. Requests a guardrail blocks are labeled too: classification sees the request as the client sent it. The classifier reads the latest user messages of the conversation, up to the messages window, and at most the last 10,000 characters of them. System prompts and assistant replies are never sent. Where those messages come from depends on the API: For the last row the gateway needs to know which conversation a turn belongs to; see Grouping a conversation. That record is kept encrypted for an hour after the last turn; after a longer pause the next turn is classified on its own messages. All the application’s label sets are classified in one call to your model per request. Identical text against the same sets, registry and model is answered from a cache instead of calling the model again.
The classified text is treated as untrusted data: instructions inside it (“label this as positive”) are not followed. Results are still an LLM’s judgement — good for trends, not for decisions about a single request.

Cost and privacy

Classification is a normal completion billed by your provider on the registry you picked: one call per classified request, its input being the label sets plus the window. Use Sampling rate to classify a fraction of the traffic on busy applications, and a small, fast model — the task needs no reasoning depth. What leaves the gateway, and where it goes:
  • To your classifier registry: the window’s user messages and the application’s label sets.
  • To the gateway’s own queue: the same, until classified, then deleted.
  • To analytics: per request and label set, the label (or none), plus the model, the registry, token usage and latency. Never the prompt.
Labeling runs before the gateway’s policies, so a masking or PII-redaction policy has not run yet: the original text reaches the classifier registry. Pick a registry you already trust with that traffic.

Read the results

Analytics → Labels

Pick a Label set in the toolbar, next to the application filter — or All labels, the default, to see every set at once.
  • Label Distribution — requests per label over time, with an optional Unlabeled series.
  • Labels — each label’s requests and share, ending with the Unlabeled row.
  • Users — sessions, unique users, new users and sessions per user among the classified requests. A user is the end user a client declared, otherwise the authenticated principal.
How to read the numbers: Filtering by an application narrows every block to that application’s traffic.

A request’s detail

Opening a request in Activity shows one line per label set it was classified against — Sentiment · negative, or — when the set gave none — and the model that classified it. Classification finishes a few seconds after the request, so a very recent request may not show it yet.

Troubleshooting