More builders. More throughput. Better Mercury 2.

10x free tokens. 10x higher rate limits. A more capable Mercury 2.

Inception Team

Over the past few months, thousands of developers have started building with Mercury across real-time voice applications, search and retrieval pipelines, and AI coding subagents. As we’ve scaled capacity, we’ve also continued improving the model itself.

Today, we’re making Mercury easier to build with.

100M free tokens, up from 10M

Every new Inception API key now includes 100 million free tokens. Enough headroom to benchmark Mercury against your current stack on production workloads.

Start building
10x higher rate limits

We've increased free-tier rate limits by 10x, so you can run production-like traffic without hitting a wall.

A faster, more capable Mercury 2

Since launch, we’ve continued improving Mercury 2 across production workloads:

  • Lower time-to-first-token

  • Lower end-to-end latency

  • More reliable tool calling

These improvements are especially noticeable for real-time voice, search, coding agents, and multi-agent workflows, where latency compounds across every model call.

Today, dozens of AI-native companies and enterprises run Mercury 2 in production.

Now available on Baseten

Mercury 2 is live on Baseten today as part of the launch of Baseten for Model Labs. If your team already builds there, you can add Mercury 2 to your stack without onboarding a new provider.

Start building

For enterprise rate limits, tighter latency budgets, SLAs, or help tuning a specific workload, contact hello@inceptionlabs.ai. Response within an hour.


The future of LLMs is here

The future of LLMs is here