Voice Agents
The real-time reasoning LLM, built for interactive applications. 1,000 tokens/sec on standard NVIDIA GPUs.
Hear the latency disappear
A live agent, reasoning and tool calls in real time. Start the demo and turn up the volume.
Quality at full speed
Tool-calling vs. Latency
Time to first token, p50 (log scale) → slower
Cost
roughly half a cent per minute of conversation
Latency
e2e latency with instant mode
Quality
on IFBench vs GPT 4.1
on Tau3Bench vs GPT 4.1
In voice, latency isn't a backend metric. It's the experience. Every extra pause makes an agent feel less capable and less trustworthy.
Traditional LLMs generate one token at a time, too slow for live conversation. Mercury generates in parallel, so it reasons in real time without the wait.
Built for the thinking layer of voice
Real-time reasoning, built for voice agents. The intelligence and speed a live call demands.

Customer support
Resolve hard calls on the first try, with reasoning fast enough to feel human.

Front desk
Pick up every call instantly and actually handle it, at a cost that scales to all of them.

Patient care
Supercharge editorial and creative work—less waiting, more creating.

Education
Instantly surface the right data from across your organization’s knowledge base.

Character game
Characters that reason in the moment, with no lag to break the immersion.






