Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Yuxian, Hu, Qinghao, Yang, Shang, Xi, Haocheng, Chen, Junyu, Han, Song, Cai, Han
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915518773985280
author Gu, Yuxian
Hu, Qinghao
Yang, Shang
Xi, Haocheng
Chen, Junyu
Han, Song
Cai, Han
author_facet Gu, Yuxian
Hu, Qinghao
Yang, Shang
Xi, Haocheng
Chen, Junyu
Han, Song
Cai, Han
contents We present Jet-Nemotron, a new family of hybrid-architecture language models, which matches or exceeds the accuracy of leading full-attention models while significantly improving generation throughput. Jet-Nemotron is developed using Post Neural Architecture Search (PostNAS), a novel neural architecture exploration pipeline that enables efficient model design. Unlike prior approaches, PostNAS begins with a pre-trained full-attention model and freezes its MLP weights, allowing efficient exploration of attention block designs. The pipeline includes four key components: (1) learning optimal full-attention layer placement and elimination, (2) linear attention block selection, (3) designing new attention blocks, and (4) performing hardware-aware hyperparameter search. Our Jet-Nemotron-2B model achieves comparable or superior accuracy to Qwen3, Qwen2.5, Gemma3, and Llama3.2 across a comprehensive suite of benchmarks while delivering up to 53.6x generation throughput speedup and 6.1x prefilling speedup. It also achieves higher accuracy on MMLU and MMLU-Pro than recent advanced MoE full-attention models, such as DeepSeek-V3-Small and Moonlight, despite their larger scale with 15B total and 2.2B activated parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15884
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
Gu, Yuxian
Hu, Qinghao
Yang, Shang
Xi, Haocheng
Chen, Junyu
Han, Song
Cai, Han
Computation and Language
Artificial Intelligence
Machine Learning
We present Jet-Nemotron, a new family of hybrid-architecture language models, which matches or exceeds the accuracy of leading full-attention models while significantly improving generation throughput. Jet-Nemotron is developed using Post Neural Architecture Search (PostNAS), a novel neural architecture exploration pipeline that enables efficient model design. Unlike prior approaches, PostNAS begins with a pre-trained full-attention model and freezes its MLP weights, allowing efficient exploration of attention block designs. The pipeline includes four key components: (1) learning optimal full-attention layer placement and elimination, (2) linear attention block selection, (3) designing new attention blocks, and (4) performing hardware-aware hyperparameter search. Our Jet-Nemotron-2B model achieves comparable or superior accuracy to Qwen3, Qwen2.5, Gemma3, and Llama3.2 across a comprehensive suite of benchmarks while delivering up to 53.6x generation throughput speedup and 6.1x prefilling speedup. It also achieves higher accuracy on MMLU and MMLU-Pro than recent advanced MoE full-attention models, such as DeepSeek-V3-Small and Moonlight, despite their larger scale with 15B total and 2.2B activated parameters.
title Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.15884