Efficient Stagewise Pretraining via Progressive Subnetworks
Fuente:
arXiv
Saved in:
| Main Authors: | Panigrahi, Abhishek, Saunshi, Nikunj, Lyu, Kaifeng, Miryoosefi, Sobhan, Reddi, Sashank, Kale, Satyen, Kumar, Sanjiv |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Landscape-Aware Growing: The Power of a Little LAG
by: Karp, Stefani, et al.
Published: (2024)
by: Karp, Stefani, et al.
Published: (2024)
On the Inductive Bias of Stacking Towards Improving Reasoning
by: Saunshi, Nikunj, et al.
Published: (2024)
by: Saunshi, Nikunj, et al.
Published: (2024)
Reasoning with Latent Thoughts: On the Power of Looped Transformers
by: Saunshi, Nikunj, et al.
Published: (2025)
by: Saunshi, Nikunj, et al.
Published: (2025)
On the Role of Depth and Looping for In-Context Learning with Task Diversity
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
by: Rawat, Ankit Singh, et al.
Published: (2024)
by: Rawat, Ankit Singh, et al.
Published: (2024)
AdaBoN: Adaptive Best-of-N Alignment
by: Raman, Vinod, et al.
Published: (2025)
by: Raman, Vinod, et al.
Published: (2025)
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models
by: Bhaskar, Adithya, et al.
Published: (2024)
by: Bhaskar, Adithya, et al.
Published: (2024)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
by: Bayazit, Deniz, et al.
Published: (2023)
by: Bayazit, Deniz, et al.
Published: (2023)
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
Structured Preconditioners in Adaptive Optimization: A Unified Analysis
by: Xie, Shuo, et al.
Published: (2025)
by: Xie, Shuo, et al.
Published: (2025)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
by: Zhou, Yongchao, et al.
Published: (2023)
by: Zhou, Yongchao, et al.
Published: (2023)
Loss Landscape Degeneracy and Stagewise Development in Transformers
by: Hoogland, Jesse, et al.
Published: (2024)
by: Hoogland, Jesse, et al.
Published: (2024)
Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe
by: You, Chong, et al.
Published: (2025)
by: You, Chong, et al.
Published: (2025)
RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
by: Wen, Kaiyue, et al.
Published: (2024)
by: Wen, Kaiyue, et al.
Published: (2024)
Asynchronous Local-SGD Training for Language Modeling
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
Representing Rule-based Chatbots with Transformers
by: Friedman, Dan, et al.
Published: (2024)
by: Friedman, Dan, et al.
Published: (2024)
Faster Rates For Federated Variational Inequalities
by: Wang, Guanghui, et al.
Published: (2026)
by: Wang, Guanghui, et al.
Published: (2026)
HierRouter: Coordinated Routing of Specialized Large Language Models via Reinforcement Learning
by: Gupta, Nikunj, et al.
Published: (2025)
by: Gupta, Nikunj, et al.
Published: (2025)
Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
by: Jin, Tian, et al.
Published: (2025)
by: Jin, Tian, et al.
Published: (2025)
How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
by: Malladi, Sadhika, et al.
Published: (2022)
by: Malladi, Sadhika, et al.
Published: (2022)
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
by: Panigrahi, Abhishek, et al.
Published: (2025)
by: Panigrahi, Abhishek, et al.
Published: (2025)
Towards Efficient Active Learning in NLP via Pretrained Representations
by: Vysogorets, Artem, et al.
Published: (2024)
by: Vysogorets, Artem, et al.
Published: (2024)
Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
by: Gu, Xinran, et al.
Published: (2025)
by: Gu, Xinran, et al.
Published: (2025)
SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection
by: Tang, Kexian, et al.
Published: (2026)
by: Tang, Kexian, et al.
Published: (2026)
Group-Level Data Selection for Efficient Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
by: Krishnakumar, Arjun, et al.
Published: (2025)
by: Krishnakumar, Arjun, et al.
Published: (2025)
Eager Updates For Overlapped Communication and Computation in DiLoCo
by: Kale, Satyen, et al.
Published: (2025)
by: Kale, Satyen, et al.
Published: (2025)
Mimetic Initialization Helps State Space Models Learn to Recall
by: Trockman, Asher, et al.
Published: (2024)
by: Trockman, Asher, et al.
Published: (2024)
On Importance of Pruning and Distillation for Efficient Low Resource NLP
by: Mirashi, Aishwarya, et al.
Published: (2024)
by: Mirashi, Aishwarya, et al.
Published: (2024)
The Power of Power Law: Asymmetry Enables Compositional Reasoning
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
Deep sequence models tend to memorize geometrically; it is unclear why
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
Training Large Reasoning Models Efficiently via Progressive Thought Encoding
by: Zhang, Zeliang, et al.
Published: (2026)
by: Zhang, Zeliang, et al.
Published: (2026)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
by: Ma, Xuezhe, et al.
Published: (2024)
by: Ma, Xuezhe, et al.
Published: (2024)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
Are More Tokens Rational? Inference-Time Scaling in Language Models as Adaptive Resource Rationality
by: Hu, Zhimin, et al.
Published: (2026)
by: Hu, Zhimin, et al.
Published: (2026)
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
by: Chen, Tianhao, et al.
Published: (2025)
by: Chen, Tianhao, et al.
Published: (2025)
Efficient and Flexible Topic Modeling using Pretrained Embeddings and Bag of Sentences
by: Schneider, Johannes
Published: (2023)
by: Schneider, Johannes
Published: (2023)
Similar Items
-
Landscape-Aware Growing: The Power of a Little LAG
by: Karp, Stefani, et al.
Published: (2024) -
On the Inductive Bias of Stacking Towards Improving Reasoning
by: Saunshi, Nikunj, et al.
Published: (2024) -
Reasoning with Latent Thoughts: On the Power of Looped Transformers
by: Saunshi, Nikunj, et al.
Published: (2025) -
On the Role of Depth and Looping for In-Context Learning with Task Diversity
by: Gatmiry, Khashayar, et al.
Published: (2024) -
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
by: Gatmiry, Khashayar, et al.
Published: (2024)