Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Akter, Syeda Nahida, Prabhumoye, Shrimai, Nyberg, Eric, Patwary, Mostofa, Shoeybi, Mohammad, Choi, Yejin, Catanzaro, Bryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
RLP: Reinforcement as a Pretraining Objective
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2025)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2025)
Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
von: Feng, Steven, et al.
Veröffentlicht: (2024)
von: Feng, Steven, et al.
Veröffentlicht: (2024)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
von: Mahabadi, Rabeeh Karimi, et al.
Veröffentlicht: (2025)
von: Mahabadi, Rabeeh Karimi, et al.
Veröffentlicht: (2025)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
von: Lu, Ximing, et al.
Veröffentlicht: (2025)
von: Lu, Ximing, et al.
Veröffentlicht: (2025)
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
von: Jung, Jaehun, et al.
Veröffentlicht: (2025)
von: Jung, Jaehun, et al.
Veröffentlicht: (2025)
Data, Data Everywhere: A Guide for Pretraining Dataset Construction
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
von: Cui, Brandon, et al.
Veröffentlicht: (2026)
von: Cui, Brandon, et al.
Veröffentlicht: (2026)
VISREAS: Complex Visual Reasoning with Unanswerable Questions
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
Self-Imagine: Effective Unimodal Reasoning with Multimodal Models using Self-Imagination
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
von: Su, Dan, et al.
Veröffentlicht: (2024)
von: Su, Dan, et al.
Veröffentlicht: (2024)
Compact Language Models via Pruning and Knowledge Distillation
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024)
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024)
iGRPO: Self-Feedback-Driven LLM Reasoning
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining
von: Wang, Boxin, et al.
Veröffentlicht: (2023)
von: Wang, Boxin, et al.
Veröffentlicht: (2023)
Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression
von: Tasnim, Nazia, et al.
Veröffentlicht: (2026)
von: Tasnim, Nazia, et al.
Veröffentlicht: (2026)
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
von: Lee, Chankyu, et al.
Veröffentlicht: (2024)
von: Lee, Chankyu, et al.
Veröffentlicht: (2024)
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
von: Xu, Chejian, et al.
Veröffentlicht: (2025)
von: Xu, Chejian, et al.
Veröffentlicht: (2025)
Nemotron-4 15B Technical Report
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
RAVEN: In-Context Learning with Retrieval-Augmented Encoder-Decoder Language Models
von: Huang, Jie, et al.
Veröffentlicht: (2023)
von: Huang, Jie, et al.
Veröffentlicht: (2023)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities
von: Xu, Peng, et al.
Veröffentlicht: (2024)
von: Xu, Peng, et al.
Veröffentlicht: (2024)
ChatQA: Surpassing GPT-4 on Conversational QA and RAG
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
AgentKit: Structured LLM Reasoning with Dynamic Graphs
von: Wu, Yue, et al.
Veröffentlicht: (2024)
von: Wu, Yue, et al.
Veröffentlicht: (2024)
RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs
von: Yu, Yue, et al.
Veröffentlicht: (2024)
von: Yu, Yue, et al.
Veröffentlicht: (2024)
On Data Engineering for Scaling LLM Terminal Capabilities
von: Pi, Renjie, et al.
Veröffentlicht: (2026)
von: Pi, Renjie, et al.
Veröffentlicht: (2026)
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
von: Yang, Zhuolin, et al.
Veröffentlicht: (2026)
von: Yang, Zhuolin, et al.
Veröffentlicht: (2026)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2025)
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2025)
Upcycling Large Language Models into Mixture of Experts
von: He, Ethan, et al.
Veröffentlicht: (2024)
von: He, Ethan, et al.
Veröffentlicht: (2024)
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
LLM Pruning and Distillation in Practice: The Minitron Approach
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2024)
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2024)
Multi-Agent Evolve: LLM Self-Improve through Co-evolution
von: Chen, Yixing, et al.
Veröffentlicht: (2025)
von: Chen, Yixing, et al.
Veröffentlicht: (2025)
LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
von: Elango, Venmugil, et al.
Veröffentlicht: (2026)
von: Elango, Venmugil, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024) -
RLP: Reinforcement as a Pretraining Objective
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2025) -
Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025) -
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
von: Feng, Steven, et al.
Veröffentlicht: (2024) -
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
von: Mahabadi, Rabeeh Karimi, et al.
Veröffentlicht: (2025)