/v1/audio/speech) takes JSON and returns raw audio bytes with the
upstream content type — not a JSON wrapper, so pipe it straight to a file.
Transcriptions (/v1/audio/transcriptions) is multipart with model and
file.
/v1/audio/translations is not served.
Which registries can serve it
OpenAI, Azure OpenAI, OpenRouter, Groq, Mistral and OpenAI-compatible endpoints do both speech and transcription. The gateway smooths the one real difference: Mistral wantsvoice_id and returns base64 in
JSON; a client sends OpenAI’s voice and receives raw bytes regardless.
xAI’s voice API is a different wire and not covered. Anthropic, Gemini, Vertex,
Bedrock, Cohere, DeepSeek and Cerebras have no audio routes.
How it routes
Speech and transcription are separate capabilities. A registry can have one without the other, and the pool is narrowed per request to the one being asked for. From there routing applies; pinning an incapable registry is a400, an empty pool a 503. Multipart bodies are forwarded
unchanged.