

Firas Trabelsi

Yanis Miraoui

Samar Khanna

Xinyu Zhao

Gokul Gunasekaran

Emily Liu

Kenan Hasanaliyev
Two weeks ago, alongside Mercury 2.5, we previewed Mercury Voice. Today it's generally available for enterprise customers.
Mercury Voice is a diffusion LLM (dLLM) tuned to power voice agents. It reasons, calls tools, and follows long system prompts while keeping latency low enough for natural conversation. On a test set of real customer-service prompts, Mercury Voice returns its first answer token in under 320 milliseconds (median), while also beating models like GPT-6 Luna and Gemma 4 31B on a suite of agentic and voice benchmarks.
At a glance
Speed: 320 ms median (p50) time to first answer token (p95: 750 ms) on production voice prompts, 5.9x faster than GPT-6 Luna (no reasoning).
Quality: Outperforms models including Gemma 4 31B, GPT-6 Luna, GLM-5.3-Flash, and Qwen3.5-397B on a composite of agentic and conversational benchmarks.
Reasoning effort: Three settings (low, medium, high)
Context: 128K tokens, with up to 50K output tokens.
Price: $0.40 per million input tokens and $1.50 per million output tokens
At launch, Mercury Voice is 50% off, at $0.20 per million input and $0.75 per million output.
Low-latency conversation
For voice applications to feel natural, the agent has to respond within about 500 ms of the caller finishing; any longer and the pause feels awkward. That's why we optimize for time to first answer token (TTFAT): how long until the model starts producing the words the caller actually hears, after all reasoning is done.

Mercury Voice at low effort is the only model with a median (p50) TTFAT under the 500 ms budget, and its p95 (750 ms) is faster than the median of every model except GPT-OSS-120B (low) on Cerebras. In voice applications, tail latency matters almost as much as the median, since a p95 of four seconds means roughly one turn in twenty fails.
Reasoning in real-time
Voice application builders have traditionally had to choose between an intelligent, reasoning model that is slow or a non-reasoning model that can only follow basic instructions but is fast. Due to tight latency budgets, non-reasoning models like Gemma 4 31B (no reasoning) and GPT-4.1 are the default for voice agents.
Mercury Voice removes the trade-off. Here's how it compares:

Indeed, Mercury Voice outperforms newer models, such as GPT-6 Luna, GLM-5.3-Flash, and Gemini 3.5 Flash-Lite, on a composite benchmark that includes τ³-bench Telecom, τ³-bench Retail, τ³-bench Airline, IFBench, and BFCL v4.
The difference is even more noticeable when quality and latency are plotted together. When averaged across multiple conversational agent benchmarks such as, τ³-bench Telecom, τ³-bench Retail, and τ³-bench Airline, Mercury Voice achieves higher quality than almost every other model we tested, and is twice as fast as the next fastest model.

Benchmarks only go so far in voice. Here's what 320 milliseconds sounds like in a live conversation.
Mercury Voice in production
Enterprises are already using Mercury Voice to power intelligent, low-latency voice applications.
Audivi AI uses Mercury Voice to power fully automated drive-thru ordering, handling real-time customer conversations, complex order modifications, and targeted upsells without a human in the loop.
At the drive-thru, the order is won or lost in the pause. Mercury Voice shortened ours without making us dumb down the agent. Once you hear the difference, it's hard to accept anything slower.
Jason Riggs, Chief Product Officer, Audivi AI
Altur builds autonomous voice agents that handle calls for financial institutions, negotiating payment plans and handling objections in real time. In Altur's internal benchmarks, Mercury Voice was the fastest model tested by a clear margin.
Mercury Voice is by far the fastest model we've tested. It's the first time the language model isn't the bottleneck in our voice pipeline.
Pedro Fernández, Co-Founder & CTO, Altur
OpenCall builds AI phone agents that handle live customer calls. On its production workload, Mercury brought median model response latency close to 170 milliseconds.
Any company working with real-time agents has had to choose between a model that responds intelligently and one that responds instantly. You don't have to trade intelligence for speed. With Mercury, you can have both.
Oliver Silverstein, Co-founder and CEO, OpenCall
Pricing
Mercury Voice is priced at $0.40 per million input tokens and $1.50 per million output tokens. At launch, it's 50% off: $0.20 per million input and $0.75 per million output.
For a typical voice-agent profile, that works out to about $0.009 per minute of conversation. That's ~5x cheaper than GPT-4.1 ($0.045 per min).
Get started
For enterprise customers, Mercury Voice is available today through the Inception API as an OpenAI API compatible endpoint. So, it drops into the LLM slot of whatever voice agent platforms or frameworks you use: LiveKit, Pipecat, Vapi, Retell, or your own.
Enterprise customers interested in Mercury Voice can contact us at sales@inceptionlabs.ai to get access.






