MarkTechPost→ original

OpenAI Released GPT-Realtime-2.1 and Mini Version for Voice Agents in API

OpenAI released two new models in the Realtime lineup to its API—GPT-Realtime-2.1 and a lightweight GPT-Realtime-2.1-mini for voice agents. The mini version works as a compact reasoning model for voice and costs the same as the previous gpt-realtime-mini. Thanks to improved caching, p95 latency has been reduced by at least 25%; connection works via the WebRTC protocol.

AI-processed from MarkTechPost; edited by Hamidun News
OpenAI Released GPT-Realtime-2.1 and Mini Version for Voice Agents in API
Source: MarkTechPost. Collage: Hamidun News.
◐ Listen to article

OpenAI added two new models to its Realtime API lineup — GPT-Realtime-2.1 and the lightweight GPT-Realtime-2.1-mini, designed for voice agents with low response latency.

What's New in the Realtime Lineup

GPT-Realtime-2.1-mini is a compact reasoning model oriented toward voice scenarios. In terms of price, it corresponds to the earlier gpt-realtime-mini, meaning the transition to the new version does not increase the cost of use for developers already working with the previous mini-model. The Realtime API lineup at OpenAI was originally created for tasks where text models with intermediate response generation are not suitable: voice assistants, conversational support interfaces, and agents that need to respond to user input with almost no pause.

The release of two models at once — the more powerful GPT-Realtime-2.1 and the lightweight mini-version — gives developers a choice between a more powerful reasoning model and a more affordable option for scenarios where speed and cost matter more than the depth of reasoning in each response.

  • Two new models: GPT-Realtime-2.1 and GPT-Realtime-2.1-mini
  • GPT-Realtime-2.1-mini — compact reasoning model for voice
  • Price of mini-version comparable to previous gpt-realtime-mini
  • p95 latency reduced by at least 25% through improved caching
  • Connection to models — via WebRTC protocol

Why p95 Latency Matters

For voice agents that must respond to users in real time, latency is one of the key quality indicators: the higher it is, the more mechanical and unnatural the conversation feels. The p95 metric shows the latency that is not exceeded by 95% of requests, meaning it reflects the behavior of the system not in the best case but in the typical and near-worst case. A reduction of this metric by at least 25% through improved caching means that the vast majority of voice responses the system now delivers significantly faster than before, without changing the reasoning model itself.

How to Connect to the New Models

OpenAI preserved for the Realtime lineup the connection via the WebRTC protocol — the same method used for previous versions. This allows developers who have already integrated voice agents based on the Realtime API to switch to the new models without changing the transport protocol.

What This Means

The Realtime lineup update shows that OpenAI continues to make targeted improvements to voice agent infrastructure — reducing latency and making reasoning models for voice more affordable, rather than only releasing more powerful but also more expensive versions.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Need AI working inside your business — not just in your newsfeed?

I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).

What do you think?
Loading comments…