A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
Fuente:
arXiv
Saved in:
| Main Authors: | Rawat, Ankit Singh, Sadhanala, Veeranjaneyulu, Rostamizadeh, Afshin, Chakrabarti, Ayan, Jitkrittum, Wittawat, Feinberg, Vladimir, Kim, Seungyeon, Harutyunyan, Hrayr, Saunshi, Nikunj, Nado, Zachary, Shivanna, Rakesh, Reddi, Sashank J., Menon, Aditya Krishna, Anil, Rohan, Kumar, Sanjiv |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Landscape-Aware Growing: The Power of a Little LAG
by: Karp, Stefani, et al.
Published: (2024)
by: Karp, Stefani, et al.
Published: (2024)
Reasoning with Latent Thoughts: On the Power of Looped Transformers
by: Saunshi, Nikunj, et al.
Published: (2025)
by: Saunshi, Nikunj, et al.
Published: (2025)
On the Role of Depth and Looping for In-Context Learning with Task Diversity
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
On the Inductive Bias of Stacking Towards Improving Reasoning
by: Saunshi, Nikunj, et al.
Published: (2024)
by: Saunshi, Nikunj, et al.
Published: (2024)
Faster Cascades via Speculative Decoding
by: Narasimhan, Harikrishna, et al.
Published: (2024)
by: Narasimhan, Harikrishna, et al.
Published: (2024)
Efficient Stagewise Pretraining via Progressive Subnetworks
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
When Does Confidence-Based Cascade Deferral Suffice?
by: Jitkrittum, Wittawat, et al.
Published: (2023)
by: Jitkrittum, Wittawat, et al.
Published: (2023)
Language Model Cascades: Token-level uncertainty and beyond
by: Gupta, Neha, et al.
Published: (2024)
by: Gupta, Neha, et al.
Published: (2024)
Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation
by: Lukasik, Michal, et al.
Published: (2025)
by: Lukasik, Michal, et al.
Published: (2025)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
by: Zhou, Yongchao, et al.
Published: (2023)
by: Zhou, Yongchao, et al.
Published: (2023)
Mimetic Initialization Helps State Space Models Learn to Recall
by: Trockman, Asher, et al.
Published: (2024)
by: Trockman, Asher, et al.
Published: (2024)
Exponential Family Trend Filtering on Lattices
by: Sadhanala, Veeranjaneyulu, et al.
Published: (2022)
by: Sadhanala, Veeranjaneyulu, et al.
Published: (2022)
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
by: Bae, Sangmin, et al.
Published: (2024)
by: Bae, Sangmin, et al.
Published: (2024)
SpacTor-T5: Pre-training T5 Models with Span Corruption and Replaced Token Detection
by: Ye, Ke, et al.
Published: (2024)
by: Ye, Ke, et al.
Published: (2024)
Multivariate Trend Filtering for Lattice Data
by: Sadhanala, Veeranjaneyulu, et al.
Published: (2021)
by: Sadhanala, Veeranjaneyulu, et al.
Published: (2021)
Cascade-Aware Training of Language Models
by: Wang, Congchao, et al.
Published: (2024)
by: Wang, Congchao, et al.
Published: (2024)
Structured Preconditioners in Adaptive Optimization: A Unified Analysis
by: Xie, Shuo, et al.
Published: (2025)
by: Xie, Shuo, et al.
Published: (2025)
SoftSRV: Learn to Generate Targeted Synthetic Data
by: DeSalvo, Giulia, et al.
Published: (2024)
by: DeSalvo, Giulia, et al.
Published: (2024)
Algorithms for Learning Kernels Based on Centered Alignment
by: Cortes, Corinna, et al.
Published: (2012)
by: Cortes, Corinna, et al.
Published: (2012)
Efficient Document Ranking with Learnable Late Interactions
by: Ji, Ziwei, et al.
Published: (2024)
by: Ji, Ziwei, et al.
Published: (2024)
Universal Model Routing for Efficient LLM Inference
by: Jitkrittum, Wittawat, et al.
Published: (2025)
by: Jitkrittum, Wittawat, et al.
Published: (2025)
In-context Learning in Presence of Spurious Correlations
by: Harutyunyan, Hrayr, et al.
Published: (2024)
by: Harutyunyan, Hrayr, et al.
Published: (2024)
Continuous Chain of Thought Enables Parallel Exploration and Reasoning
by: Gozeten, Halil Alperen, et al.
Published: (2025)
by: Gozeten, Halil Alperen, et al.
Published: (2025)
Cost-Aware Routing for Efficient Text-To-Image Generation
by: Li, Qinchan, et al.
Published: (2025)
by: Li, Qinchan, et al.
Published: (2025)
Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs
by: Yang, Xiulin, et al.
Published: (2025)
by: Yang, Xiulin, et al.
Published: (2025)
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel
by: Cutler, Dylan, et al.
Published: (2025)
by: Cutler, Dylan, et al.
Published: (2025)
A Little Aggression Goes a Long Way
by: Krishnan, Jyothi, et al.
Published: (2024)
by: Krishnan, Jyothi, et al.
Published: (2024)
A Little Confidence Goes a Long Way
by: Scoville, John, et al.
Published: (2024)
by: Scoville, John, et al.
Published: (2024)
Simplicity Bias via Global Convergence of Sharpness Minimization
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Dynamic Exponent Market Maker: Personalized Portfolio Manager and One Pool to Trade Them All
by: Kositwattanarerk, Wittawat
Published: (2025)
by: Kositwattanarerk, Wittawat
Published: (2025)
Rethinking FID: Towards a Better Evaluation Metric for Image Generation
by: Jayasumana, Sadeep, et al.
Published: (2023)
by: Jayasumana, Sadeep, et al.
Published: (2023)
Gatekeeper: Improving Model Cascades Through Confidence Tuning
by: Rabanser, Stephan, et al.
Published: (2025)
by: Rabanser, Stephan, et al.
Published: (2025)
A Little Human Data Goes A Long Way
by: Ashok, Dhananjay, et al.
Published: (2024)
by: Ashok, Dhananjay, et al.
Published: (2024)
EXAMINING LOCAL SOCIAL IDENTITIES THROUGH PATTERNS OF BIOLOGICAL AND CULTURAL VARIATION IN THE SOLCOR AYLLU, SAN PEDRO DE ATACAMA, CHILE
by: Kristin L. Nado
Published: (2012)
by: Kristin L. Nado
Published: (2012)
Experimentos en una ciencia no experimental
by: Hrayr Der Hagopian Tlapanco
Published: (2016)
by: Hrayr Der Hagopian Tlapanco
Published: (2016)
Think before you speak: Training Language Models With Pause Tokens
by: Goyal, Sachin, et al.
Published: (2023)
by: Goyal, Sachin, et al.
Published: (2023)
Swarming behaviour associated with group cohesion in tree-dwelling bats
by: Naďo, Ladislav, et al.
Published: (2015)
by: Naďo, Ladislav, et al.
Published: (2015)
Sequential Fair Allocation With Replenishments: A Little Envy Goes An Exponentially Long Way
by: Onyeze, Chido, et al.
Published: (2025)
by: Onyeze, Chido, et al.
Published: (2025)
Similar Items
-
Landscape-Aware Growing: The Power of a Little LAG
by: Karp, Stefani, et al.
Published: (2024) -
Reasoning with Latent Thoughts: On the Power of Looped Transformers
by: Saunshi, Nikunj, et al.
Published: (2025) -
On the Role of Depth and Looping for In-Context Learning with Task Diversity
by: Gatmiry, Khashayar, et al.
Published: (2024) -
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
by: Gatmiry, Khashayar, et al.
Published: (2024) -
On the Inductive Bias of Stacking Towards Improving Reasoning
by: Saunshi, Nikunj, et al.
Published: (2024)