Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
Fuente:
arXiv
Guardado en:
| Autores principales: | Fan, Zhiyuan, Liu, Yifeng, Zhao, Qingyue, Yuan, Angela, Gu, Quanquan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
por: Chen, Zixiang, et al.
Publicado: (2025)
por: Chen, Zixiang, et al.
Publicado: (2025)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
por: Zhao, Qingyue, et al.
Publicado: (2026)
por: Zhao, Qingyue, et al.
Publicado: (2026)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
por: Zhao, Qingyue, et al.
Publicado: (2025)
por: Zhao, Qingyue, et al.
Publicado: (2025)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
por: Ji, Kaixuan, et al.
Publicado: (2026)
por: Ji, Kaixuan, et al.
Publicado: (2026)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
por: Ji, Kaixuan, et al.
Publicado: (2026)
por: Ji, Kaixuan, et al.
Publicado: (2026)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
por: He, Jiafan, et al.
Publicado: (2025)
por: He, Jiafan, et al.
Publicado: (2025)
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
por: Chen, Zixiang, et al.
Publicado: (2024)
por: Chen, Zixiang, et al.
Publicado: (2024)
Deep Delta Learning
por: Zhang, Yifan, et al.
Publicado: (2026)
por: Zhang, Yifan, et al.
Publicado: (2026)
Accelerated Preference Optimization for Large Language Model Alignment
por: He, Jiafan, et al.
Publicado: (2024)
por: He, Jiafan, et al.
Publicado: (2024)
Tensor Product Attention Is All You Need
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
Survival Models: Proper Scoring Rule and Stochastic Optimization with Competing Risks
por: Alberge, Julie, et al.
Publicado: (2024)
por: Alberge, Julie, et al.
Publicado: (2024)
A Universal Model for Human Mobility Prediction
por: Long, Qingyue, et al.
Publicado: (2024)
por: Long, Qingyue, et al.
Publicado: (2024)
CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing
por: Yang, Chen, et al.
Publicado: (2024)
por: Yang, Chen, et al.
Publicado: (2024)
MARS-M: When Variance Reduction Meets Matrices
por: Liu, Yifeng, et al.
Publicado: (2025)
por: Liu, Yifeng, et al.
Publicado: (2025)
Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation
por: Yuan, Huizhuo, et al.
Publicado: (2024)
por: Yuan, Huizhuo, et al.
Publicado: (2024)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
por: He, Di, et al.
Publicado: (2025)
por: He, Di, et al.
Publicado: (2025)
OSDTW: Optimal Shared Depth and Task Weighting for Long-Tailed Recognition
por: Chu, Chang, et al.
Publicado: (2026)
por: Chu, Chang, et al.
Publicado: (2026)
Exploring Layerwise Adversarial Robustness Through the Lens of t-SNE
por: Valentim, Inês, et al.
Publicado: (2024)
por: Valentim, Inês, et al.
Publicado: (2024)
Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMs
por: Gao, Chang, et al.
Publicado: (2025)
por: Gao, Chang, et al.
Publicado: (2025)
Distributional Regression with Tabular Foundation Models: Evaluating Probabilistic Predictions via Proper Scoring Rules
por: Landsgesell, Jonas, et al.
Publicado: (2026)
por: Landsgesell, Jonas, et al.
Publicado: (2026)
Layerwise LQR for Geometry-Aware Optimization of Deep Networks
por: Dufort-Labbé, Simon, et al.
Publicado: (2026)
por: Dufort-Labbé, Simon, et al.
Publicado: (2026)
The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws
por: Zaher, Eslam, et al.
Publicado: (2026)
por: Zaher, Eslam, et al.
Publicado: (2026)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
por: Pan, Rui, et al.
Publicado: (2024)
por: Pan, Rui, et al.
Publicado: (2024)
Group Representational Position Encoding
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
Transformers Trained via Gradient Descent Can Provably Learn a Class of Teacher Models
por: Zhang, Chenyang, et al.
Publicado: (2026)
por: Zhang, Chenyang, et al.
Publicado: (2026)
Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-Finetuning
por: Ye, Jiasheng, et al.
Publicado: (2023)
por: Ye, Jiasheng, et al.
Publicado: (2023)
Understanding Transferable Representation Learning and Zero-shot Transfer in CLIP
por: Chen, Zixiang, et al.
Publicado: (2023)
por: Chen, Zixiang, et al.
Publicado: (2023)
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
por: Zhang, Junkai, et al.
Publicado: (2024)
por: Zhang, Junkai, et al.
Publicado: (2024)
Outlier-weighed Layerwise Sampling for LLM Fine-tuning
por: Li, Pengxiang, et al.
Publicado: (2024)
por: Li, Pengxiang, et al.
Publicado: (2024)
RSPO: Regularized Self-Play Alignment of Large Language Models
por: Tang, Xiaohang, et al.
Publicado: (2025)
por: Tang, Xiaohang, et al.
Publicado: (2025)
Fast Sampling via Discrete Non-Markov Diffusion Models with Predetermined Transition Time
por: Chen, Zixiang, et al.
Publicado: (2023)
por: Chen, Zixiang, et al.
Publicado: (2023)
Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification
por: Lutz, Patrick, et al.
Publicado: (2026)
por: Lutz, Patrick, et al.
Publicado: (2026)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
por: Ye, Chenlu, et al.
Publicado: (2026)
por: Ye, Chenlu, et al.
Publicado: (2026)
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
por: He, Di, et al.
Publicado: (2026)
por: He, Di, et al.
Publicado: (2026)
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
por: Zhao, Heyang, et al.
Publicado: (2025)
por: Zhao, Heyang, et al.
Publicado: (2025)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
por: Bergsma, Shane, et al.
Publicado: (2025)
por: Bergsma, Shane, et al.
Publicado: (2025)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
por: Lei, Ge, et al.
Publicado: (2025)
por: Lei, Ge, et al.
Publicado: (2025)
OLC-WA: Drift Aware Tuning-Free Online Classification with Weighted Average
por: Shaira, Mohammad Abu, et al.
Publicado: (2025)
por: Shaira, Mohammad Abu, et al.
Publicado: (2025)
Employing Layerwised Unsupervised Learning to Lessen Data and Loss Requirements in Forward-Forward Algorithms
por: Hwang, Taewook, et al.
Publicado: (2024)
por: Hwang, Taewook, et al.
Publicado: (2024)
Ejemplares similares
-
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
por: Chen, Zixiang, et al.
Publicado: (2025) -
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
por: Zhao, Qingyue, et al.
Publicado: (2026) -
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
por: Zhao, Qingyue, et al.
Publicado: (2025) -
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
por: Ji, Kaixuan, et al.
Publicado: (2026) -
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
por: Ji, Kaixuan, et al.
Publicado: (2026)