Nemotron-4 15B Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | Parmar, Jupinder, Prabhumoye, Shrimai, Jennings, Joseph, Patwary, Mostofa, Subramanian, Sandeep, Su, Dan, Zhu, Chen, Narayanan, Deepak, Jhunjhunwala, Aastha, Dattagupta, Ayush, Jawa, Vibhu, Liu, Jiwei, Mahabaleshwarkar, Ameya, Nitski, Osvald, Brundyn, Annika, Maki, James, Martinez, Miguel, You, Jiaxuan, Kamalu, John, LeGresley, Patrick, Fridman, Denys, Casper, Jared, Aithal, Ashwath, Kuchaiev, Oleksii, Shoeybi, Mohammad, Cohen, Jonathan, Catanzaro, Bryan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data, Data Everywhere: A Guide for Pretraining Dataset Construction
by: Parmar, Jupinder, et al.
Published: (2024)
by: Parmar, Jupinder, et al.
Published: (2024)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
by: Akter, Syeda Nahida, et al.
Published: (2025)
by: Akter, Syeda Nahida, et al.
Published: (2025)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
by: Feng, Steven, et al.
Published: (2024)
by: Feng, Steven, et al.
Published: (2024)
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
by: Taghibakhshi, Ali, et al.
Published: (2025)
by: Taghibakhshi, Ali, et al.
Published: (2025)
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
by: Parmar, Jupinder, et al.
Published: (2024)
by: Parmar, Jupinder, et al.
Published: (2024)
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
by: Akter, Syeda Nahida, et al.
Published: (2025)
by: Akter, Syeda Nahida, et al.
Published: (2025)
Nemotron-4 340B Technical Report
by: Nvidia, et al.
Published: (2024)
by: Nvidia, et al.
Published: (2024)
RLP: Reinforcement as a Pretraining Objective
by: Hatamizadeh, Ali, et al.
Published: (2025)
by: Hatamizadeh, Ali, et al.
Published: (2025)
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
by: Su, Dan, et al.
Published: (2024)
by: Su, Dan, et al.
Published: (2024)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
by: Lu, Ximing, et al.
Published: (2025)
by: Lu, Ximing, et al.
Published: (2025)
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
by: Jung, Jaehun, et al.
Published: (2025)
by: Jung, Jaehun, et al.
Published: (2025)
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
by: Taghibakhshi, Ali, et al.
Published: (2025)
by: Taghibakhshi, Ali, et al.
Published: (2025)
LLM Pruning and Distillation in Practice: The Minitron Approach
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2024)
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2024)
Student gender modulates the intersection of calculus proficiency and calculus self-efficacy in an introductory electricity and magnetism course
by: Fischer, Christopher J., et al.
Published: (2024)
by: Fischer, Christopher J., et al.
Published: (2024)
FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
by: Liu, Zihan, et al.
Published: (2025)
by: Liu, Zihan, et al.
Published: (2025)
Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression
by: Tasnim, Nazia, et al.
Published: (2026)
by: Tasnim, Nazia, et al.
Published: (2026)
Compact Language Models via Pruning and Knowledge Distillation
by: Muralidharan, Saurav, et al.
Published: (2024)
by: Muralidharan, Saurav, et al.
Published: (2024)
Upcycling Large Language Models into Mixture of Experts
by: He, Ethan, et al.
Published: (2024)
by: He, Ethan, et al.
Published: (2024)
Llama-Nemotron: Efficient Reasoning Models
by: Bercovich, Akhiad, et al.
Published: (2025)
by: Bercovich, Akhiad, et al.
Published: (2025)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
by: Wang, Boxin, et al.
Published: (2025)
by: Wang, Boxin, et al.
Published: (2025)
When2Call: When (not) to Call Tools
by: Ross, Hayley, et al.
Published: (2025)
by: Ross, Hayley, et al.
Published: (2025)
NVIDIA Nemotron Nano V2 VL
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
by: Diao, Shizhe, et al.
Published: (2025)
by: Diao, Shizhe, et al.
Published: (2025)
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
nach0: Multimodal Natural and Chemical Languages Foundation Model
by: Livne, Micha, et al.
Published: (2023)
by: Livne, Micha, et al.
Published: (2023)
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
by: Yang, Zhuolin, et al.
Published: (2026)
by: Yang, Zhuolin, et al.
Published: (2026)
The AI Consumer Index (ACE)
by: Benchek, Julien, et al.
Published: (2025)
by: Benchek, Julien, et al.
Published: (2025)
NVIDIA Nemotron 3: Efficient and Open Intelligence
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Corrupción entre particulares: lesividad de la conducta y consecuencias en sede de tipificación de acuerdo al análisis comparado
by: Osvald Artaza Varela
Published: (2019)
by: Osvald Artaza Varela
Published: (2019)
NVIDIA Nemotron Parse 1.1
by: Chumachenko, Kateryna, et al.
Published: (2025)
by: Chumachenko, Kateryna, et al.
Published: (2025)
Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
by: Zhang, Shaokun, et al.
Published: (2025)
by: Zhang, Shaokun, et al.
Published: (2025)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
by: Liu, Zihan, et al.
Published: (2024)
by: Liu, Zihan, et al.
Published: (2024)
iGRPO: Self-Feedback-Driven LLM Reasoning
by: Hatamizadeh, Ali, et al.
Published: (2026)
by: Hatamizadeh, Ali, et al.
Published: (2026)
Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
by: NVIDIA, et al.
Published: (2026)
by: NVIDIA, et al.
Published: (2026)
Marginal Fermi liquids from Fermi surfaces coupled via matrix boson gas
by: Mishra, Vibhu
Published: (2025)
by: Mishra, Vibhu
Published: (2025)
Similar Items
-
Data, Data Everywhere: A Guide for Pretraining Dataset Construction
by: Parmar, Jupinder, et al.
Published: (2024) -
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025) -
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
by: Akter, Syeda Nahida, et al.
Published: (2024) -
Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
by: Akter, Syeda Nahida, et al.
Published: (2025) -
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
by: Feng, Steven, et al.
Published: (2024)