One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | He, Di, Tu, Songjun, Wang, Keyu, Yin, Lu, Liu, Shiwei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
by: He, Di, et al.
Published: (2025)
by: He, Di, et al.
Published: (2025)
Outlier-weighed Layerwise Sampling for LLM Fine-tuning
by: Li, Pengxiang, et al.
Published: (2024)
by: Li, Pengxiang, et al.
Published: (2024)
Infinite Width Models That Work: Why Feature Learning Doesn't Matter as Much as You Think
by: Sernau, Luke
Published: (2024)
by: Sernau, Luke
Published: (2024)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
by: He, Qianxi, et al.
Published: (2025)
by: He, Qianxi, et al.
Published: (2025)
Bayesian Elicitation with LLMs: Model Size Helps, Extra "Reasoning" Doesn't Always
by: Hobor, Luka, et al.
Published: (2026)
by: Hobor, Luka, et al.
Published: (2026)
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
by: Falahati, Ali, et al.
Published: (2026)
by: Falahati, Ali, et al.
Published: (2026)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
by: Cooper, A. Feder, et al.
Published: (2024)
by: Cooper, A. Feder, et al.
Published: (2024)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
by: Wang, Keyu, et al.
Published: (2025)
by: Wang, Keyu, et al.
Published: (2025)
MedCalc-Bench Doesn't Measure What You Think: A Benchmark Audit and the Case for Open-Book Evaluation
by: Krohn-Grimberghe, Artus
Published: (2026)
by: Krohn-Grimberghe, Artus
Published: (2026)
Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies
by: Liu, Ming
Published: (2026)
by: Liu, Ming
Published: (2026)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
by: Lei, Ge, et al.
Published: (2025)
by: Lei, Ge, et al.
Published: (2025)
Online Preference-based Reinforcement Learning with Self-augmented Feedback from Large Language Model
by: Tu, Songjun, et al.
Published: (2024)
by: Tu, Songjun, et al.
Published: (2024)
ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
by: Li, Victoria R., et al.
Published: (2024)
by: Li, Victoria R., et al.
Published: (2024)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
by: Lu, Liming, et al.
Published: (2026)
by: Lu, Liming, et al.
Published: (2026)
Revisiting Intermediate-Layer Matching in Knowledge Distillation: Layer-Selection Strategy Doesn't Matter (Much)
by: Yu, Zony, et al.
Published: (2025)
by: Yu, Zony, et al.
Published: (2025)
Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation
by: Yuan, Peiwen, et al.
Published: (2025)
by: Yuan, Peiwen, et al.
Published: (2025)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
by: Yin, Lu, et al.
Published: (2023)
by: Yin, Lu, et al.
Published: (2023)
Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMs
by: Gao, Chang, et al.
Published: (2025)
by: Gao, Chang, et al.
Published: (2025)
One-Topic-Doesn't-Fit-All: Transcreating Reading Comprehension Test for Personalized Learning
by: Han, Jieun, et al.
Published: (2025)
by: Han, Jieun, et al.
Published: (2025)
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
by: Li, Pengxiang, et al.
Published: (2024)
by: Li, Pengxiang, et al.
Published: (2024)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
by: Verma, Gaurav, et al.
Published: (2024)
by: Verma, Gaurav, et al.
Published: (2024)
Beyond One-Size-Fits-All: Adaptive Subgraph Denoising for Zero-Shot Graph Learning with Large Language Models
by: Li, Fengzhi, et al.
Published: (2026)
by: Li, Fengzhi, et al.
Published: (2026)
In-Dataset Trajectory Return Regularization for Offline Preference-based Reinforcement Learning
by: Tu, Songjun, et al.
Published: (2024)
by: Tu, Songjun, et al.
Published: (2024)
Beyond One-Size-Fits-All: Adapting Counterfactual Explanations to User Objectives
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2024)
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2024)
No One Size Fits All: QueryBandits for Hallucination Mitigation
by: Cho, Nicole, et al.
Published: (2026)
by: Cho, Nicole, et al.
Published: (2026)
One Fits All: General Mobility Trajectory Modeling via Masked Conditional Diffusion
by: Long, Qingyue, et al.
Published: (2025)
by: Long, Qingyue, et al.
Published: (2025)
Adaptive Heavy-Tailed Stochastic Gradient Descent
by: Gong, Bodu, et al.
Published: (2025)
by: Gong, Bodu, et al.
Published: (2025)
Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't
by: Taguchi, Chihiro, et al.
Published: (2024)
by: Taguchi, Chihiro, et al.
Published: (2024)
HTMuon: Improving Muon via Heavy-Tailed Spectral Correction
by: Pang, Tianyu, et al.
Published: (2026)
by: Pang, Tianyu, et al.
Published: (2026)
TailedTS: Benchmark Dataset for Heavy-Tailed Time Series Prediction and Periodicity Quantification
by: Chen, Xinyu, et al.
Published: (2026)
by: Chen, Xinyu, et al.
Published: (2026)
One Size Doesn't Fit All: Age-Aware Gamification Mechanics for Multimedia Learning Environments
by: Kaißer, Sarah, et al.
Published: (2025)
by: Kaißer, Sarah, et al.
Published: (2025)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
by: Fan, Zhiyuan, et al.
Published: (2025)
by: Fan, Zhiyuan, et al.
Published: (2025)
Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
by: Dang, Quy-Anh, et al.
Published: (2025)
by: Dang, Quy-Anh, et al.
Published: (2025)
Robust Offline Reinforcement learning with Heavy-Tailed Rewards
by: Zhu, Jin, et al.
Published: (2023)
by: Zhu, Jin, et al.
Published: (2023)
One Jump Is All You Need: Short-Cutting Transformers for Early Exit Prediction with One Jump to Fit All Exit Levels
by: Seshadri, Amrit Diggavi
Published: (2025)
by: Seshadri, Amrit Diggavi
Published: (2025)
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
by: Wang, Ziqiao, et al.
Published: (2025)
by: Wang, Ziqiao, et al.
Published: (2025)
$(ε, u)$-Adaptive Regret Minimization in Heavy-Tailed Bandits
by: Genalti, Gianmarco, et al.
Published: (2023)
by: Genalti, Gianmarco, et al.
Published: (2023)
Layerwise LQR for Geometry-Aware Optimization of Deep Networks
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
From Compression to Expression: A Layerwise Analysis of In-Context Learning
by: Jiang, Jiachen, et al.
Published: (2025)
by: Jiang, Jiachen, et al.
Published: (2025)
Similar Items
-
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
by: He, Di, et al.
Published: (2025) -
Outlier-weighed Layerwise Sampling for LLM Fine-tuning
by: Li, Pengxiang, et al.
Published: (2024) -
Infinite Width Models That Work: Why Feature Learning Doesn't Matter as Much as You Think
by: Sernau, Luke
Published: (2024) -
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
by: He, Qianxi, et al.
Published: (2025) -
Bayesian Elicitation with LLMs: Model Size Helps, Extra "Reasoning" Doesn't Always
by: Hobor, Luka, et al.
Published: (2026)