Mercury 2.5
Recorded: Sept. 8, 2026, 9:09 p.m.
| Original | Summarized |
Introducing Mercury 2.5 – Inception Introducing Mercury 2.5|Read the blogIntroducing Mercury 2.5ModelsEnterpriseCompanyResearchTry APIContact SalesModelsEnterpriseCompanyResearchTry APIContact SalesBlog/ProductIntroducing Mercury 2.5More intelligence at Mercury speedStefano ErmonCEOToday, we’re releasing Mercury 2.5, our most capable production model yet. It is a significant step up in quality over Mercury 2, with the same low-latency, low-cost serving profile. Since Mercury 2’s launch, thousands of developers have built with it, dozens of enterprises have put it into production, and usage has grown over an order of magnitude. It now serves latency-sensitive workloads across search, voice, and coding products.Those workloads gave us a clearer signal than benchmarks alone. We used customer feedback and production failure cases to sharpen the evals and focus training. Mercury 2.5 is the first result of that loop.What changedMercury 2.5 is the most capable diffusion LLM on the market. To our knowledge, it is the largest diffusion language model ever trained.Quality: 40% increase in intelligence from Mercury 2. Comparable to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Speed: 1,107 tokens per second on widely-available NVIDIA GPUs.Context: 260K tokens.Price: $0.20 per million input and $0.75 per million output.At launch, Mercury 2.5 is 80% off at $0.04 per million input and $0.15 per million output.Capabilities: Tunable reasoning, parallel tool calls, and schema-aligned JSON.Since Mercury 2's launch, we've watched Inception advance diffusion-based language models further on NVIDIA AI infrastructure. Mercury 2.5's step up in intelligence paired with sustained speeds and low costs, reflects how quickly new architectures can mature into production-ready systems on the NVIDIA platform.Shruti Koparkar, Senior Manager of Product, Accelerated Computing Group at NVIDIAMercury in productionSearch Agents and RAG pipelinesOne search request can trigger dozens of model calls: plan the search, rewrite queries, rerank results, structure facts, summarize sources, and check the answer. Mercury keeps those calls fast enough to stay inside a single user interaction. Several leading search-infrastructure companies now run it in production.Voice agents and interactive applicationsIn voice, latency isn’t an infrastructure detail. It is the pause a caller hears.OpenCall builds AI phone agents that handle live customer calls. On its production workload, Mercury brought median model response latency close to 170 milliseconds.After we switched to Mercury, our P99 response time dropped from several minutes to just one second, and our P50 dropped from 0.4 seconds to under 0.2 — significantly faster than any other provider we’ve seen, and that’s including reasoning.Oliver Silverstein, Co-founder and CEO, OpenCallRead more: The first reasoning model fast enough to pick up the phone Coding subagents and assistantsCoding agents already split work across models. One may plan or write code while others search, run tools, route requests, summarize state, or compact a long session. Those supporting calls happen again and again, so latency and cost compound quickly.Augment Code uses Mercury for context compaction, model routing, and MCP tool search. Moving compaction to Mercury cut latency by 82%, from roughly 150 seconds to 27 seconds, and reduced cost by 90% while maintaining quality. Tool-search summaries return in under a second.The same speed applies to developing web apps. Watch Mercury 2.5 generate a working music discovery log web app from a few prompts in the demo below.Mercury Voice and Mercury Router PreviewAlongside Mercury 2.5, we’re announcing a preview of Mercury Voice and Mercury Router. Mercury Voice delivers time-to-first-token (TTFT) under 170 milliseconds and is a dLLM optimized for voice agents with the tightest latency budgets. Mercury Router understands incoming prompts with a dLLM and routes them to the best models (open and closed models) that offer the best mix of quality, speed, and cost.Get StartedMercury models are available through our Inception API, Baseten, and OpenRouter. Enterprise deployments support dedicated capacity, autoscaling, compliance controls, and configurable data retention.Try Mercury 2.5 in chat Try the API with 100 million free tokens · Read the API docsBaseten customers: Deploy Mercury 2.5 through your existing Baseten setup.Y Combinator companies: Claim $500,000 in deployment benefits.Evaluating Mercury for voice? We’ll work with you to test workload fit, and validate performance under your serving constraints. Contact us.What’s NextWe have already started training our next model. It is our largest model yet, and we are targeting a release in the coming months. Our next model will be a leap in capability without giving up diffusion’s speed and token-efficiency. That requires progress on model training, inference, evals, and infrastructure. If that's the kind of problem you want to work on, we’d love to hear from you.More soon.Product·Sep 8, 2026Introducing Mercury 2.5Product·Sep 8, 2026Introducing Mercury 2.5Product·Aug 11, 2026Mercury 2 for Search: Fast enough to run a hundred times per queryProduct·Aug 11, 2026Mercury 2 for Search: Fast enough to run a hundred times per queryProduct·Jul 29, 2026More builders. More throughput. Better Mercury 2.Product·Jul 29, 2026More builders. More throughput. Better Mercury 2.Product·Sep 8, 2026Introducing Mercury 2.5Product·Aug 11, 2026Mercury 2 for Search: Fast enough to run a hundred times per queryProduct·Jul 29, 2026More builders. More throughput. Better Mercury 2.The future of LLMs is hereGet StartedThe future of LLMs is hereGet StartedProductsGet StartedModelsPricingCompanyAbout UsResearchCareersBlogResourcesMercury ChatAPI PlatformDocumentationIntegrationsPartnersLegalTerms of ServicePrivacy PolicyCookie SettingsContactSalesInquiresDiscordXLinkedIn© 2026 InceptionProductsGet StartedModelsPricingCompanyAbout UsResearchCareersBlogResourcesMercury ChatAPI PlatformDocumentationIntegrationsPartnersLegalTerms of ServicePrivacy PolicyCookie SettingsContactSalesInquiresDiscordXLinkedIn© 2026 Inception |
Mercury 2.5 is presented as the most capable production model developed by Inception, representing a significant advancement in quality over its predecessor, Mercury 2, while crucially maintaining low-latency and low-cost serving profiles. This evolution was driven by incorporating customer feedback and analysis of production failure cases to hone the training and evaluation process, establishing it as the first result of this iterative improvement loop. Technically, Mercury 2.5 is positioned as the largest diffusion language model ever trained, boasting a forty percent increase in intelligence, making it comparable to cost-optimized frontier models such as GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Its operational specifications include a processing speed of one thousand one hundred seven tokens per second on widely available NVIDIA GPUs, a context window of twenty-six thousand tokens, and pricing structured at $0.20 per million input and $0.75 per million output, with launch pricing offering substantial discounts. Furthermore, Mercury 2.5 incorporates advanced capabilities including tunable reasoning, parallel tool calls, and schema-aligned JSON generation. The utility of Mercury 2.5 is most evident in its ability to handle latency-sensitive workloads across domains like search, voice agents, and coding. In the realm of search agents and Retrieval Augmented Generation pipelines, the model allows a single search request to trigger complex sequences involving planning, query rewriting, result reranking, fact structuring, and source summarization, all while maintaining rapid processing speeds within a single user interaction. This efficiency has been demonstrated in voice applications where latency is paramount; for instance, OpenCall utilized Mercury to reduce median model response latency for live customer calls to nearly one hundred seventy milliseconds, achieving significant reductions in P99 response time from several minutes to one second and P50 from 0.4 seconds to under 0.2 seconds. In coding environments, such as Augmented Code, Mercury was used for context compaction, routing requests, and tool searching, which resulted in an eighty-two percent reduction in latency and a ninety percent reduction in cost while preserving quality. In addition to core model capabilities, Inception is previewing related services, including Mercury Voice, which is optimized for voice agents with a time-to-first-token under one hundred seventy milliseconds, and Mercury Router, designed to intelligently route incoming prompts to the optimal models based on their quality, speed, and cost metrics. These models are accessible through the Inception API, Baseten, and OpenRouter, offering enterprise deployment options that include dedicated capacity management and compliance controls. The development effort continues with plans for a larger subsequent model, aiming to achieve further leaps in capability while sustaining diffusion architecture's inherent speed and token efficiency. |