Inception Launches Mercury 2.5, the Next Tier of Intelligence for Diffusion LLMs

Inception Launches Mercury 2.5, the Next Tier of Intelligence for Diffusion LLMs
View on original source
Category: SciTech
Share
Archive
Like
REDWOOD CITY, Calif.--(BUSINESS WIRE)--Sep 8, 2026-- Inception, the company behind the first commercial diffusion large language models (dLLMs), today announced the launch of Mercury 2.5, the most capable dLLM and the fastest reasoning LLM in production. Mercury 2.5 delivers the next tier of intelligence for dLLMs and runs over 1,100 tokens per second in production. This press release features multimedia. View the full release here: https://www.businesswire.com/news/home/20260908593295/en/ Most LLMs in production today still generate text autoregressively: one token at a time. That approach ties cost and latency directly to reasoning depth, so the more a model has to think, the slower and more expensive it becomes. dLLMs work differently, starting with a rough draft of the output and refining tokens in parallel. dLLMs have moved from a research bet to a production workhorse. Enterprise Mercury usage is growing by over an order of magnitude, powering search, voice, and coding agents at scale. Inception proved that diffusion was enterprise-ready with Mercury 2, matching Claude Haiku and GPT Mini intelligence at roughly 10x the throughput, and Mercury models have since become the dLLM of choice for teams that can't afford to trade intelligence for speed, or speed for cost. Mercury 2.5 builds directly on that foundation: more intelligence, lower cost, and the same speedup over traditional models. 'Nobody in this industry thinks LLMs can get smarter, faster, and cheaper at the same time. It's a limitation of autoregressive modeling,' said Stefano Ermon, CEO and co-founder of Inception. 'Mercury 2.5 proves the switch to diffusion opens that door.' Token Efficiency As AI infrastructure costs climb and data centers strain to keep up with demand, token efficiency is under more scrutiny than it was even a year ago. Every token costs compute, power, and data center capacity, and autoregressive models spend all three generating one token at a time. dLLMs make fundamentally better use of each GPU cycle, generating many tokens per step instead of one, which means more intelligence per dollar. That efficiency runs on widely-available NVIDIA GPUs, not specialized silicon. Smarter, Not Slower Mercury 2.5 jumps 10 points in intelligence over Mercury 2, making it more intelligent than any previous dLLM, comparable to cost-optimized frontier models like GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Throughput is up to over 1,100 tokens per second, and pricing drops to $0.20 / $0.75 per 1M tokens. At launch, Mercury 2.5 is 80% off at $0.04 / $0.15 per 1M tokens. Mercury 2.5 extends the context window to 260K tokens, up from 128K, and it adds tunable reasoning, native tool use, JSON mode, and parallel tool calls. 'After switching to Mercury, our P99 response time dropped from several minutes to just one second, and P50 dropped from 0.4 seconds to less than 0.2 seconds, faster than any other provider we've seen on the market, including reasoning,' said Oliver Silverstein, Co-founder and CEO of OpenCall. Where Diffusion Thrives Mercury dLLMs are already running in production across the workloads that demand responsive inference: Alongside Mercury 2.5, Inception also announced a preview of Mercury Voice and Mercury Router. Mercury Voice delivers time-to-first-token (TTFT) under 170ms and is a dLLM optimized for voice agents with the tightest latency budgets. Mercury Router understands incoming prompts with a dLLM and routes them to the best models (open and closed models) that offer the best mix of quality, speed, and cost. Interested customers can contact sales@inceptionlabs.ai to get preview access to Mercury Voice and Mercury Router. Mercury 2.5 is enterprise-ready and available now on Inception's API, OpenRouter, and Baseten. Get started with 100 million free tokens: https://platform.inceptionlabs.ai/ About Inception Inception develops diffusion-based large language models (dLLMs) designed for efficient, low-latency AI applications. While traditional autoregressive LLMs generate text sequentially, Inception's diffusion-based models generate outputs in parallel, enabling faster inference and improved reliability for real-time use cases like search, voice, and coding agents. Inception's Mercury models are available via the Inception API, OpenRouter, and Baseten. Based in Redwood City, California, Inception is backed by Menlo Ventures, Mayfield, Innovation Endeavors, M12 (Microsoft's venture capital fund), Snowflake Ventures, Databricks Ventures, and individual backers including Andrew Ng, Andrej Karpathy, and Eric Schmidt. For more information, visit www.inceptionlabs.ai.

(0)Comments

 

A note on cookies

Newshunt uses essential cookies to keep you signed in and to remember your language and country, so the site works the way you expect. With your permission, we'd also like to use analytics cookies to understand how people use Newshunt and improve it over time.

Accepting only affects analytics. To learn more, view our Privacy Policy or Terms & Conditions.