Mercury: Ultra-Fast Language Models Based on Diffusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Labs, Inception, Khanna, Samar, Kharbanda, Siddhant, Li, Shufan, Varma, Harshit, Wang, Eric, Birnbaum, Sawyer, Luo, Ziyang, Miraoui, Yanis, Palrecha, Akash, Ermon, Stefano, Grover, Aditya, Kuleshov, Volodymyr
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918066670010368
author Labs, Inception
Khanna, Samar
Kharbanda, Siddhant
Li, Shufan
Varma, Harshit
Wang, Eric
Birnbaum, Sawyer
Luo, Ziyang
Miraoui, Yanis
Palrecha, Akash
Ermon, Stefano
Grover, Aditya
Kuleshov, Volodymyr
author_facet Labs, Inception
Khanna, Samar
Kharbanda, Siddhant
Li, Shufan
Varma, Harshit
Wang, Eric
Birnbaum, Sawyer
Luo, Ziyang
Miraoui, Yanis
Palrecha, Akash
Ermon, Stefano
Grover, Aditya
Kuleshov, Volodymyr
contents We present Mercury, a new generation of commercial-scale large language models (LLMs) based on diffusion. These models are parameterized via the Transformer architecture and trained to predict multiple tokens in parallel. In this report, we detail Mercury Coder, our first set of diffusion LLMs designed for coding applications. Currently, Mercury Coder comes in two sizes: Mini and Small. These models set a new state-of-the-art on the speed-quality frontier. Based on independent evaluations conducted by Artificial Analysis, Mercury Coder Mini and Mercury Coder Small achieve state-of-the-art throughputs of 1109 tokens/sec and 737 tokens/sec, respectively, on NVIDIA H100 GPUs and outperform speed-optimized frontier models by up to 10x on average while maintaining comparable quality. We discuss additional results on a variety of code benchmarks spanning multiple languages and use-cases as well as real-world validation by developers on Copilot Arena, where the model currently ranks second on quality and is the fastest model overall. We also release a public API at https://platform.inceptionlabs.ai/ and free playground at https://chat.inceptionlabs.ai
format Preprint
id arxiv_https___arxiv_org_abs_2506_17298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mercury: Ultra-Fast Language Models Based on Diffusion
Labs, Inception
Khanna, Samar
Kharbanda, Siddhant
Li, Shufan
Varma, Harshit
Wang, Eric
Birnbaum, Sawyer
Luo, Ziyang
Miraoui, Yanis
Palrecha, Akash
Ermon, Stefano
Grover, Aditya
Kuleshov, Volodymyr
Computation and Language
Artificial Intelligence
Machine Learning
We present Mercury, a new generation of commercial-scale large language models (LLMs) based on diffusion. These models are parameterized via the Transformer architecture and trained to predict multiple tokens in parallel. In this report, we detail Mercury Coder, our first set of diffusion LLMs designed for coding applications. Currently, Mercury Coder comes in two sizes: Mini and Small. These models set a new state-of-the-art on the speed-quality frontier. Based on independent evaluations conducted by Artificial Analysis, Mercury Coder Mini and Mercury Coder Small achieve state-of-the-art throughputs of 1109 tokens/sec and 737 tokens/sec, respectively, on NVIDIA H100 GPUs and outperform speed-optimized frontier models by up to 10x on average while maintaining comparable quality. We discuss additional results on a variety of code benchmarks spanning multiple languages and use-cases as well as real-world validation by developers on Copilot Arena, where the model currently ranks second on quality and is the fastest model overall. We also release a public API at https://platform.inceptionlabs.ai/ and free playground at https://chat.inceptionlabs.ai
title Mercury: Ultra-Fast Language Models Based on Diffusion
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.17298