Voice Agents

Deploy real-time voice agents

Deploy real-time voice agents

Deploy real-time voice agents

The real-time reasoning LLM, built for interactive applications. 1,000 tokens/sec on standard NVIDIA GPUs.

Trusted by leading enterprises

Trusted by leading enterprises

Hear the latency disappear

A live agent, reasoning and tool calls in real time. Start the demo and turn up the volume.

Waiting for VoiceChat components…

Quality at full speed

All other LLMs generate text one token at a time. Mercury diffusion LLMs (dLLMs) generate tokens in parallel, increasing speed.

Tool-calling vs. Latency

Time to first token, p50 (log scale) → slower

All other LLMs generate text one token at a time. Mercury diffusion LLMs (dLLMs) generate tokens in parallel, increasing speed.

Cost

~$0.005/min

~$0.005/min

~$0.005/min

roughly half a cent per minute of conversation

Latency

170ms

170ms

170ms

e2e latency with instant mode

Quality

+27

+27

+27

on IFBench vs GPT 4.1

+24

+24

+24

on Tau3Bench vs GPT 4.1

At OpenCall, we have been using Mercury 2 to power our production voice agents handling complex patient calls. In our testing, Mercury 2 outperformed GPT OSS 120B on Cerebras on instruction-following, tool-use and reasoning through multi-step workflows. It gives us the reasoning quality we need without sacrificing the latency required for a natural phone call experience.
Oliver Silverstein, CEO
At OpenCall, we have been using Mercury 2 to power our production voice agents handling complex patient calls. In our testing, Mercury 2 outperformed GPT OSS 120B on Cerebras on instruction-following, tool-use and reasoning through multi-step workflows. It gives us the reasoning quality we need without sacrificing the latency required for a natural phone call experience.
Oliver Silverstein, CEO
At OpenCall, we have been using Mercury 2 to power our production voice agents handling complex patient calls. In our testing, Mercury 2 outperformed GPT OSS 120B on Cerebras on instruction-following, tool-use and reasoning through multi-step workflows. It gives us the reasoning quality we need without sacrificing the latency required for a natural phone call experience.
Oliver Silverstein, CEO

In voice, latency isn't a backend metric. It's the experience. Every extra pause makes an agent feel less capable and less trustworthy.


Traditional LLMs generate one token at a time, too slow for live conversation. Mercury generates in parallel, so it reasons in real time without the wait.

Built for the thinking layer of voice

Real-time reasoning, built for voice agents. The intelligence and speed a live call demands.

Customer support

Resolve hard calls on the first try, with reasoning fast enough to feel human.

Front desk

Pick up every call instantly and actually handle it, at a cost that scales to all of them.

Patient care

Supercharge editorial and creative work—less waiting, more creating.

Education

Instantly surface the right data from across your organization’s knowledge base.

Character game

Characters that reason in the moment, with no lag to break the immersion.

The future of LLMs is here

The future of LLMs is here