/v1/batches, with a Batches page in the dashboardSubmit chat completions, embeddings, responses, or Anthropic Messages requests as one provider batch, and results stream back as NDJSON. The Batches page sits next to Logs, filters by status and date, and opens a detail drawer with each batch's progress, cost, errors, and metadata. Read more in the batch inference docs.
Seven models are new in the catalog, including two text-to-speech models billed by the second.
X-Merge-Billed-Audio-Seconds headerRead more in the Gateway model catalog docs.
A customer's Configuration tab sets its own prompt injection and data loss prevention policy, which inherits the organization policy until overridden and is enforced at request time. A new Controls tab on projects and customers overrides zero data retention, vendor restrictions, and region restrictions, and each unit resets to the organization setting on its own. Read more in the per-customer restrictions docs.
priority_orderReasoning effort takes a policy-wide default, minimum, and maximum, plus per-model overrides, all enforced at dispatch and factored into vendor choice. Pass priority_order on /v1/responses or any compatible surface to route through an ordered model list, and Gateway saves it as a visible routing policy. Read more in the routing policy docs.
/v1/responses breaks down latency per requestTransfer, gateway, and provider time come back in the response body and the streaming done chunk. Non-streaming calls also carry a Server-Timing header. Read more in the Responses API docs.
/v1/live/sessionsEach session is backed by gpt-live-1 by OpenAI. Sessions run with zero data retention. Read more in the live voice docs.
Improvements