AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | He, Di, Tu, Songjun, Jaiswal, Ajay, Shen, Li, Yuan, Ganzhao, Liu, Shiwei, Yin, Lu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
by: He, Di, et al.
Published: (2026)
by: He, Di, et al.
Published: (2026)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
by: Yin, Lu, et al.
Published: (2023)
by: Yin, Lu, et al.
Published: (2023)
Dynamic Topic Evolution with Temporal Decay and Attention in Large Language Models
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
by: Liu, Yanjiang, et al.
Published: (2026)
by: Liu, Yanjiang, et al.
Published: (2026)
Optimal Decay Spectra for Linear Recurrences
by: Cao, Yang
Published: (2026)
by: Cao, Yang
Published: (2026)
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
by: Lu, Haiquan, et al.
Published: (2024)
by: Lu, Haiquan, et al.
Published: (2024)
Contrastive Token Learning with Similarity Decay for Repetition Suppression in Machine Translation
by: Dai, Huangyu, et al.
Published: (2024)
by: Dai, Huangyu, et al.
Published: (2024)
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
by: Fan, Zhiyuan, et al.
Published: (2025)
by: Fan, Zhiyuan, et al.
Published: (2025)
AlphaAgent: LLM-Driven Alpha Mining with Regularized Exploration to Counteract Alpha Decay
by: Tang, Ziyi, et al.
Published: (2025)
by: Tang, Ziyi, et al.
Published: (2025)
Oblivion: Self-Adaptive Agentic Memory Control through Decay-Driven Activation
by: Rana, Ashish, et al.
Published: (2026)
by: Rana, Ashish, et al.
Published: (2026)
LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning
by: Liu, Zihang, et al.
Published: (2025)
by: Liu, Zihang, et al.
Published: (2025)
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
by: Aswal, Darpan, et al.
Published: (2025)
by: Aswal, Darpan, et al.
Published: (2025)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
by: Jaiswal, Ajay, et al.
Published: (2024)
by: Jaiswal, Ajay, et al.
Published: (2024)
Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning
by: Bandari, Abhinav, et al.
Published: (2024)
by: Bandari, Abhinav, et al.
Published: (2024)
Outlier-weighed Layerwise Sampling for LLM Fine-tuning
by: Li, Pengxiang, et al.
Published: (2024)
by: Li, Pengxiang, et al.
Published: (2024)
Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition
by: Kim, Jiyeon, et al.
Published: (2024)
by: Kim, Jiyeon, et al.
Published: (2024)
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
by: Cheng, Wenhua, et al.
Published: (2023)
by: Cheng, Wenhua, et al.
Published: (2023)
Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding?
by: Li, Pengxiang, et al.
Published: (2026)
by: Li, Pengxiang, et al.
Published: (2026)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
by: Cao, Mingyu, et al.
Published: (2024)
by: Cao, Mingyu, et al.
Published: (2024)
Multi-Stage Balanced Distillation: Addressing Long-Tail Challenges in Sequence-Level Knowledge Distillation
by: Zhou, Yuhang, et al.
Published: (2024)
by: Zhou, Yuhang, et al.
Published: (2024)
Elucidating the Design Space of Decay in Linear Attention
by: Qin, Zhen, et al.
Published: (2025)
by: Qin, Zhen, et al.
Published: (2025)
Diffusion Language Models Know the Answer Before Decoding
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
by: Lv, Junlin, et al.
Published: (2024)
by: Lv, Junlin, et al.
Published: (2024)
The 17% Gap: Quantifying Epistemic Decay in AI-Assisted Survey Papers
by: İlter, H. Kemal
Published: (2026)
by: İlter, H. Kemal
Published: (2026)
Can We Edit LLMs for Long-Tail Biomedical Knowledge?
by: Yi, Xinhao, et al.
Published: (2025)
by: Yi, Xinhao, et al.
Published: (2025)
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
by: Chen, Yuhan, et al.
Published: (2024)
by: Chen, Yuhan, et al.
Published: (2024)
Search or Accelerate: Confidence-Switched Position Beam Search for Diffusion Language Models
by: Cao, Mingyu, et al.
Published: (2026)
by: Cao, Mingyu, et al.
Published: (2026)
MCTS-SQL: Light-Weight LLMs can Master the Text-to-SQL through Monte Carlo Tree Search
by: Yuan, Shuozhi, et al.
Published: (2025)
by: Yuan, Shuozhi, et al.
Published: (2025)
MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
by: Liu, Yuxi, et al.
Published: (2025)
by: Liu, Yuxi, et al.
Published: (2025)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
by: Pan, Birong, et al.
Published: (2025)
by: Pan, Birong, et al.
Published: (2025)
How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
by: Das, Nilanjana, et al.
Published: (2026)
by: Das, Nilanjana, et al.
Published: (2026)
AlphaAdam:Asynchronous Masked Optimization with Dynamic Alpha for Selective Updates
by: Chang, Da, et al.
Published: (2025)
by: Chang, Da, et al.
Published: (2025)
Balancing Knowledge Updates: Toward Unified Modular Editing in LLMs
by: Liu, Jiahao, et al.
Published: (2025)
by: Liu, Jiahao, et al.
Published: (2025)
eSapiens's DEREK Module: Deep Extraction & Reasoning Engine for Knowledge with LLMs
by: Shi, Isaac, et al.
Published: (2025)
by: Shi, Isaac, et al.
Published: (2025)
Controllable Spoken Dialogue Generation: An LLM-Driven Grading System for K-12 Non-Native English Learners
by: Yuan, Haidong, et al.
Published: (2026)
by: Yuan, Haidong, et al.
Published: (2026)
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
by: Fu, Yuqian, et al.
Published: (2025)
by: Fu, Yuqian, et al.
Published: (2025)
Similar Items
-
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
by: He, Di, et al.
Published: (2026) -
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
by: Yin, Lu, et al.
Published: (2023) -
Dynamic Topic Evolution with Temporal Decay and Attention in Large Language Models
by: Wu, Di, et al.
Published: (2025) -
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
by: Bergsma, Shane, et al.
Published: (2025) -
Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
by: Liu, Yanjiang, et al.
Published: (2026)