How Muon's Spectral Design Benefits Generalization: A Study on Imbalanced Data
Fuente:
arXiv
Saved in:
| Main Authors: | Vasudeva, Bhavya, Deora, Puneesh, Zhao, Yize, Sharan, Vatsal, Thrampoulidis, Christos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026)
by: Vasudeva, Bhavya, et al.
Published: (2026)
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
by: Deora, Puneesh, et al.
Published: (2025)
by: Deora, Puneesh, et al.
Published: (2025)
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
by: Behnia, Tina, et al.
Published: (2025)
by: Behnia, Tina, et al.
Published: (2025)
On the Optimization and Generalization of Multi-head Attention
by: Deora, Puneesh, et al.
Published: (2023)
by: Deora, Puneesh, et al.
Published: (2023)
Why Loss Re-weighting Works If You Stop Early: Training Dynamics of Unconstrained Features
by: Zhao, Yize, et al.
Published: (2026)
by: Zhao, Yize, et al.
Published: (2026)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Transformers Learn Low Sensitivity Functions: Investigations and Implications
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Latent Concept Disentanglement in Transformer-based Language Models
by: Hong, Guan Zhe, et al.
Published: (2025)
by: Hong, Guan Zhe, et al.
Published: (2025)
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
by: Zhao, Yize, et al.
Published: (2024)
by: Zhao, Yize, et al.
Published: (2024)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
by: Thrampoulidis, Christos
Published: (2024)
by: Thrampoulidis, Christos
Published: (2024)
Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations
by: Zhao, Yize, et al.
Published: (2025)
by: Zhao, Yize, et al.
Published: (2025)
A Spectral View of Adversarially Robust Features
by: Garg, Shivam, et al.
Published: (2018)
by: Garg, Shivam, et al.
Published: (2018)
Supervised Contrastive Representation Learning: Landscape Analysis with Unconstrained Features
by: Behnia, Tina, et al.
Published: (2024)
by: Behnia, Tina, et al.
Published: (2024)
Preconditioning Benefits of Spectral Orthogonalization in Muon
by: Ma, Jianhao, et al.
Published: (2026)
by: Ma, Jianhao, et al.
Published: (2026)
DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
by: Deng, Wenlong, et al.
Published: (2024)
by: Deng, Wenlong, et al.
Published: (2024)
Limitations on Accurate, Trusted, Human-level Reasoning
by: Panigrahy, Rina, et al.
Published: (2025)
by: Panigrahy, Rina, et al.
Published: (2025)
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
by: Ye, Qilin, et al.
Published: (2025)
by: Ye, Qilin, et al.
Published: (2025)
Improved Bounds for Swap Multicalibration and Swap Omniprediction
by: Luo, Haipeng, et al.
Published: (2025)
by: Luo, Haipeng, et al.
Published: (2025)
Optimal Multiclass U-Calibration Error and Beyond
by: Luo, Haipeng, et al.
Published: (2024)
by: Luo, Haipeng, et al.
Published: (2024)
Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated Learning
by: Deng, Wenlong, et al.
Published: (2023)
by: Deng, Wenlong, et al.
Published: (2023)
Thumb on the Scale: Optimal Loss Weighting in Last Layer Retraining
by: Stromberg, Nathan, et al.
Published: (2025)
by: Stromberg, Nathan, et al.
Published: (2025)
Memorization Capacity of Multi-Head Attention in Transformers
by: Mahdavi, Sadegh, et al.
Published: (2023)
by: Mahdavi, Sadegh, et al.
Published: (2023)
On the Inherent Privacy of Zeroth Order Projected Gradient Descent
by: Gupta, Devansh, et al.
Published: (2025)
by: Gupta, Devansh, et al.
Published: (2025)
Memory capacity of two layer neural networks with smooth activations
by: Madden, Liam, et al.
Published: (2023)
by: Madden, Liam, et al.
Published: (2023)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
by: Thrampoulidis, Christos, et al.
Published: (2025)
by: Thrampoulidis, Christos, et al.
Published: (2025)
Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods
by: Taheri, Hossein, et al.
Published: (2024)
by: Taheri, Hossein, et al.
Published: (2024)
Proper Learnability and the Role of Unlabeled Data
by: Asilis, Julian, et al.
Published: (2025)
by: Asilis, Julian, et al.
Published: (2025)
Efficient Swap Multicalibration of Elicitable Properties
by: Hu, Lunjia, et al.
Published: (2025)
by: Hu, Lunjia, et al.
Published: (2025)
Stability and Multigroup Fairness in Ranking with Uncertain Predictions
by: Devic, Siddartha, et al.
Published: (2024)
by: Devic, Siddartha, et al.
Published: (2024)
When is Multicalibration Post-Processing Necessary?
by: Hansen, Dutch, et al.
Published: (2024)
by: Hansen, Dutch, et al.
Published: (2024)
On Provable Benefits of Muon in Federated Learning
by: Zhang, Xinwen, et al.
Published: (2025)
by: Zhang, Xinwen, et al.
Published: (2025)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
by: Fu, Deqing, et al.
Published: (2026)
by: Fu, Deqing, et al.
Published: (2026)
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
by: Garrod, Connall, et al.
Published: (2025)
by: Garrod, Connall, et al.
Published: (2025)
Balanced Data, Imbalanced Spectra: Unveiling Class Disparities with Spectral Imbalance
by: Kaushik, Chiraag, et al.
Published: (2024)
by: Kaushik, Chiraag, et al.
Published: (2024)
Pre-trained Large Language Models Use Fourier Features to Compute Addition
by: Zhou, Tianyi, et al.
Published: (2024)
by: Zhou, Tianyi, et al.
Published: (2024)
Sample Amplification: Increasing Dataset Size even when Learning is Impossible
by: Axelrod, Brian, et al.
Published: (2019)
by: Axelrod, Brian, et al.
Published: (2019)
Simultaneous Swap Regret Minimization via KL-Calibration
by: Luo, Haipeng, et al.
Published: (2025)
by: Luo, Haipeng, et al.
Published: (2025)
Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning
by: Lu, Binghang, et al.
Published: (2026)
by: Lu, Binghang, et al.
Published: (2026)
Similar Items
-
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026) -
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024) -
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
by: Deora, Puneesh, et al.
Published: (2025) -
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
by: Behnia, Tina, et al.
Published: (2025) -
On the Optimization and Generalization of Multi-head Attention
by: Deora, Puneesh, et al.
Published: (2023)