Back to changelog

Batch inference through /v1/batches

Batch inference runs through /v1/batches, with a Batches page in the dashboard

Submit chat completions, embeddings, responses, or Anthropic Messages requests as one provider batch, and results stream back as NDJSON. The Batches page sits next to Logs, filters by status and date, and opens a detail drawer with each batch's progress, cost, errors, and metadata. Read more in the batch inference docs.

Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and grok-4.7 are available with zero data retention

Seven models are new in the catalog, including two text-to-speech models billed by the second.

  • Claude Opus 5.5 by Anthropic, GPT-6 Sol and GPT-6 Luna by OpenAI, and grok-4.7 by xAI, each with zero data retention
  • GLM-5.3 FlashX by Z.AI
  • Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS by Google, billed per second of generated audio and reported in the X-Merge-Billed-Audio-Seconds header

Read more in the Gateway model catalog docs.

Guardrails and data controls take per-customer and per-project overrides

A customer's Configuration tab sets its own prompt injection and data loss prevention policy, which inherits the organization policy until overridden and is enforced at request time. A new Controls tab on projects and customers overrides zero data retention, vendor restrictions, and region restrictions, and each unit resets to the organization setting on its own. Read more in the per-customer restrictions docs.

Routing policies set reasoning effort and accept an inline priority_order

Reasoning effort takes a policy-wide default, minimum, and maximum, plus per-model overrides, all enforced at dispatch and factored into vendor choice. Pass priority_order on /v1/responses or any compatible surface to route through an ordered model list, and Gateway saves it as a visible routing policy. Read more in the routing policy docs.

/v1/responses breaks down latency per request

Transfer, gateway, and provider time come back in the response body and the streaming done chunk. Non-streaming calls also carry a Server-Timing header. Read more in the Responses API docs.

Real-time voice sessions run over WebSocket at /v1/live/sessions

Each session is backed by gpt-live-1 by OpenAI. Sessions run with zero data retention. Read more in the live voice docs.

‍

Improvements

  • Azure OpenAI - Adds GPT-5.6 Sol, Luna, Terra, and GPT-6 Astra, managed or bring-your-own key
  • Pareto Inference - Added as a model vendor serving GLM-5.3 Flash with zero data retention
  • Tavily - Now available in the web search tool with zero data retention