led by Stefano Ermon

Inception profile, products and governance

Diffusion language models that refine text in parallel.

  • startup

Inception suits teams whose product makes many fast model calls in a row, such as voice agents, search pipelines and coding assistants. Its newest model trades raw intelligence for speed and low price, so keep a stronger model for the hardest reasoning steps. The team is three professor-founders and 16 open roles.

Inception is a Palo Alto AI company that builds Mercury, a family of diffusion language models. Instead of writing text one token at a time, Mercury refines many tokens in parallel, which the company says reaches more than 1,000 tokens per second on standard NVIDIA GPUs through an OpenAI-compatible API.

Founded: 2024 · HQ: Palo Alto, CA, USA · Team: 11-50 · CEO: Stefano Ermon · Funding: $50M seed (announced Nov 2025), led by Menlo Ventures

About Inception

Inception, also called Inception Labs and incorporated as Inception AI, Inc., builds diffusion-based large language models. It was founded in 2024 by three university researchers who had worked together for more than a decade: Stanford professor Stefano Ermon, who is the chief executive, UCLA professor Aditya Grover and Cornell professor Volodymyr Kuleshov. Ermon's research is on the diffusion method behind image systems such as Stable Diffusion, and the company's bet is that the same refine-in-parallel idea works for text. Inception emerged from stealth in February 2025 with the first Mercury models.

The Mercury family now has five parts. Mercury 2.5, released on 8 September 2026, is the flagship reasoning model with a 260K-token context window. Mercury 2 is the previous reasoning model and stays supported, Mercury Edit 2 handles code autocomplete and next-edit suggestions, Mercury Voice targets voice agents and is sold through sales, and Mercury Router is a preview that sends each prompt to the model with the best mix of quality, speed and cost. Everything is served through an OpenAI-compatible API, and the research page lists 10 papers, including the Mercury technical paper from June 2025.

Inception's only disclosed funding is the $50 million seed round announced in November 2025. Two investors connect to the product: Inception ships Mercury 2 on Azure Foundry, which sits next to the seed money from Microsoft's M12 fund, and its speed figures are measured on NVIDIA GPUs, the same company behind NVentures in the round.

Mercury 2.5 is sold through the Inception API, Baseten and OpenRouter, and older Mercury models reached Amazon Bedrock Marketplace and SageMaker JumpStart in August 2025. Customers Inception names include Augment Code, which moved context compaction to Mercury and reports latency down 82% (about 150 seconds to 27) at 90% lower cost, and the phone-agent company OpenCall, which reports median response latency near 170 milliseconds. The docs list voice platforms such as Vapi and Retell AI as compatible with Mercury Voice.

Inception wins on speed per dollar for short, repeated calls and loses on raw intelligence. Artificial Analysis describes Mercury 2.5 as below average in intelligence, and Inception itself positions the model against small, cost-optimized models such as GPT-5.6 Luna and Gemini 3.5 Flash-Lite, not against flagships. Rivals reach similar throughput with custom chips (Groq) or with their own diffusion releases (DiffusionGemma from Google DeepMind).

The careers page, checked on 3 October 2026, lists 16 open roles, all in the Bay Area and all in office. On data, the default API terms allow training on submitted prompts unless the user opts out in settings, while the enterprise page promises no training on customer data and configurable retention. Neither a SOC 2 report nor an ISO 27001 certificate, and no trust center, turned up when HokAI checked. HokAI files Inception under startups and lists its models by provider.

Mission

Build diffusion-based language models for production, so AI behaves more like an editor than a one-way typewriter and stays fast enough to run always-on.

Products

Inception Models on HokAI

Links

Frequently Asked Questions

How much money has Inception raised, and from whom?

Inception has disclosed one round: a $50 million seed, announced on 6 November 2025. Menlo Ventures led it, and the investor list includes Mayfield, Innovation Endeavors, Microsoft's M12, Snowflake Ventures, Databricks Investment and NVIDIA's NVentures, plus angels Andrew Ng and Andrej Karpathy. No valuation and no later round have been published by Inception so far.

What does Inception sell, and what does it cost?

Mercury 2.5, the flagship reasoning model, lists at $0.20 input and $0.75 output per million tokens, with an 80% launch discount that brings it to $0.04 and $0.15. Mercury 2 and Mercury Edit 2 (autocomplete and next-edit) both list at $0.25 input and $0.75 output. Mercury Voice is enterprise-only at $0.20 input and $0.75 output after a 50% discount, and Mercury Router is a preview with pricing on request.

Does Inception train on your data, and is it certified?

The terms of use let Inception train its models on what you submit through the API unless you switch off the 'Improve the model for everyone' option in the platform's user settings. The enterprise page describes a stricter arrangement: no training on customer data, configurable retention and caching, no-logging modes, private networking and custom legal terms. Zero-retention setups are handled case by case through support, and no SOC 2 or ISO 27001 certificate or trust center was found on 3 October 2026.

Which companies compete with Inception?

Groq is the closest rival on speed: it serves open models on custom LPU chips, while Inception trains its own diffusion models and runs them on standard GPUs. Google DeepMind competes at both ends, with Gemini 3.5 Flash-Lite as a low-cost model and DiffusionGemma as an open diffusion model. OpenAI's GPT-5.6 Luna accepts images, which Mercury 2.5 does not, so Inception wins on cost and speed per call and loses on reasoning depth and input types.

What does it take to start using Mercury?

Create an account on the Inception platform, generate an API key and send a request to the chat completions endpoint with the model name mercury-2.5. The format matches OpenAI's, so most SDKs only need a new base URL, and the first call takes a few minutes. The docs and launch post cite a one-time credit of 100 million free tokens, while one onboarding step on the models page says 10 million, so check the dashboard for the real balance.

Top Alternatives

  • Groq: Pick Inception if you want its own diffusion models on standard GPUs; pick Groq if you want an inference cloud on custom chips serving other vendors' models.
  • Google DeepMind: Pick Inception for a commercial diffusion-LLM API focused on latency; pick Google DeepMind when you want Gemini breadth or open DiffusionGemma weights.
  • OpenAI: Pick Inception for latency-bound supporting calls at low cost; pick OpenAI when answer quality and image input matter more than response time.

More AI Companies on HokAI