Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Yuxian, Hu, Qinghao, Yang, Shang, Xi, Haocheng, Chen, Junyu, Han, Song, Cai, Han |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Llama-Nemotron: Efficient Reasoning Models
by: Bercovich, Akhiad, et al.
Published: (2025)
by: Bercovich, Akhiad, et al.
Published: (2025)
Elastic Architecture Search for Efficient Language Models
by: Wang, Shang
Published: (2025)
by: Wang, Shang
Published: (2025)
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
by: Zou, Dongyun, et al.
Published: (2026)
by: Zou, Dongyun, et al.
Published: (2026)
Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
by: Hu, Qinghao, et al.
Published: (2025)
by: Hu, Qinghao, et al.
Published: (2025)
SEKI: Self-Evolution and Knowledge Inspiration based Neural Architecture Search via Large Language Models
by: Cai, Zicheng, et al.
Published: (2025)
by: Cai, Zicheng, et al.
Published: (2025)
NVIDIA Nemotron 3: Efficient and Open Intelligence
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
MiniLLM: On-Policy Distillation of Large Language Models
by: Gu, Yuxian, et al.
Published: (2023)
by: Gu, Yuxian, et al.
Published: (2023)
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
by: Yang, Zhuolin, et al.
Published: (2026)
by: Yang, Zhuolin, et al.
Published: (2026)
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
by: Zhang, Shaokun, et al.
Published: (2025)
by: Zhang, Shaokun, et al.
Published: (2025)
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
by: Wang, Boxin, et al.
Published: (2025)
by: Wang, Boxin, et al.
Published: (2025)
Efficient Streaming Language Models with Attention Sinks
by: Xiao, Guangxuan, et al.
Published: (2023)
by: Xiao, Guangxuan, et al.
Published: (2023)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
by: Xiao, Guangxuan, et al.
Published: (2022)
by: Xiao, Guangxuan, et al.
Published: (2022)
Designing Domain-Specific Large Language Models: The Critical Role of Fine-Tuning in Public Opinion Simulation
by: Lin, Haocheng
Published: (2024)
by: Lin, Haocheng
Published: (2024)
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
by: NVIDIA, et al.
Published: (2026)
by: NVIDIA, et al.
Published: (2026)
Nemotron-4 15B Technical Report
by: Parmar, Jupinder, et al.
Published: (2024)
by: Parmar, Jupinder, et al.
Published: (2024)
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
Nemotron-4 340B Technical Report
by: Nvidia, et al.
Published: (2024)
by: Nvidia, et al.
Published: (2024)
R-Capsule: Compressing High-Level Plans for Efficient Large Language Model Reasoning
by: Shan, Hongyu, et al.
Published: (2025)
by: Shan, Hongyu, et al.
Published: (2025)
A Transformer-based Neural Architecture Search Method
by: Wang, Shang, et al.
Published: (2025)
by: Wang, Shang, et al.
Published: (2025)
Following the Whispers of Values: Unraveling Neural Mechanisms Behind Value-Oriented Behaviors in LLMs
by: Hu, Ling, et al.
Published: (2025)
by: Hu, Ling, et al.
Published: (2025)
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
by: Liu, Zihan, et al.
Published: (2025)
by: Liu, Zihan, et al.
Published: (2025)
W-PCA Based Gradient-Free Proxy for Efficient Search of Lightweight Language Models
by: Wang, Shang
Published: (2025)
by: Wang, Shang
Published: (2025)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
by: Yang, Shang, et al.
Published: (2025)
by: Yang, Shang, et al.
Published: (2025)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
by: Xi, Haocheng, et al.
Published: (2024)
by: Xi, Haocheng, et al.
Published: (2024)
Sparsity Induction for Accurate Post-Training Pruning of Large Language Models
by: Jiang, Minhao, et al.
Published: (2026)
by: Jiang, Minhao, et al.
Published: (2026)
EPIC: Efficient and Parallel Inference under CFG Constraints for Diffusion Language Models
by: Jin, Hyundong, et al.
Published: (2026)
by: Jin, Hyundong, et al.
Published: (2026)
LLMatic: Neural Architecture Search via Large Language Models and Quality Diversity Optimization
by: Nasir, Muhammad U., et al.
Published: (2023)
by: Nasir, Muhammad U., et al.
Published: (2023)
Hierarchical Orthogonal Residual Spread for Precise Massive Editing in Large Language Models
by: Gu, Xiaojie, et al.
Published: (2026)
by: Gu, Xiaojie, et al.
Published: (2026)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
DIP: Dynamic In-Context Planner For Diffusion Language Models
by: Li, Yang, et al.
Published: (2026)
by: Li, Yang, et al.
Published: (2026)
FastAdaSP: Multitask-Adapted Efficient Inference for Large Speech Language Model
by: Lu, Yichen, et al.
Published: (2024)
by: Lu, Yichen, et al.
Published: (2024)
Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
by: Chen, Junyu, et al.
Published: (2024)
by: Chen, Junyu, et al.
Published: (2024)
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encoding
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
Towards Efficient Post-Training via Fourier-Driven Adapter Architectures
by: Bae, Donggyun, et al.
Published: (2025)
by: Bae, Donggyun, et al.
Published: (2025)
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
by: He, Wenkun, et al.
Published: (2025)
by: He, Wenkun, et al.
Published: (2025)
Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models
by: Jiang, Haitao, et al.
Published: (2026)
by: Jiang, Haitao, et al.
Published: (2026)
Similar Items
-
Llama-Nemotron: Efficient Reasoning Models
by: Bercovich, Akhiad, et al.
Published: (2025) -
Elastic Architecture Search for Efficient Language Models
by: Wang, Shang
Published: (2025) -
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
by: Zou, Dongyun, et al.
Published: (2026) -
Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
by: Hu, Qinghao, et al.
Published: (2025) -
SEKI: Self-Evolution and Knowledge Inspiration based Neural Architecture Search via Large Language Models
by: Cai, Zicheng, et al.
Published: (2025)