Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Ruotong, Wei, Ermin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2025)
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2025)
In-Context Reward Adaptation for Robust Preference Modeling
von: Sun, Zhenyu, et al.
Veröffentlicht: (2026)
von: Sun, Zhenyu, et al.
Veröffentlicht: (2026)
SOAP: Improving and Stabilizing Shampoo using Adam
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
4-bit Shampoo for Memory-Efficient Network Training
von: Wang, Sike, et al.
Veröffentlicht: (2024)
von: Wang, Sike, et al.
Veröffentlicht: (2024)
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2026)
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2026)
Data Whitening Improves Sparse Autoencoder Learning
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
Orthogonalized Policy Optimization:Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
Group Orthogonalized Policy Optimization:Group Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
von: Ma, Chao, et al.
Veröffentlicht: (2024)
von: Ma, Chao, et al.
Veröffentlicht: (2024)
Whitening Not Recommended for Classification Tasks in LLMs
von: Forooghi, Ali, et al.
Veröffentlicht: (2024)
von: Forooghi, Ali, et al.
Veröffentlicht: (2024)
Hessian-Free Online Certified Unlearning
von: Qiao, Xinbao, et al.
Veröffentlicht: (2024)
von: Qiao, Xinbao, et al.
Veröffentlicht: (2024)
Reservoir Subspace Injection for Online ICA under Top-n Whitening
von: Xiao, Wenjun, et al.
Veröffentlicht: (2026)
von: Xiao, Wenjun, et al.
Veröffentlicht: (2026)
FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo
von: Nam, Kyunghun, et al.
Veröffentlicht: (2026)
von: Nam, Kyunghun, et al.
Veröffentlicht: (2026)
Fisher-Orthogonal Projected Natural Gradient Descent for Continual Learning
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
von: Yao, Wei, et al.
Veröffentlicht: (2025)
von: Yao, Wei, et al.
Veröffentlicht: (2025)
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
von: Shu, Yao, et al.
Veröffentlicht: (2026)
von: Shu, Yao, et al.
Veröffentlicht: (2026)
Orthogonal Subspace Projection for Continual Machine Unlearning via SVD-Based LoRA
von: Rahulamathavan, Yogachandran, et al.
Veröffentlicht: (2026)
von: Rahulamathavan, Yogachandran, et al.
Veröffentlicht: (2026)
Approximated Orthogonal Projection Unit: Stabilizing Regression Network Training Using Natural Gradient
von: Wang, Shaoqi, et al.
Veröffentlicht: (2024)
von: Wang, Shaoqi, et al.
Veröffentlicht: (2024)
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
von: Lu, Yishun, et al.
Veröffentlicht: (2025)
von: Lu, Yishun, et al.
Veröffentlicht: (2025)
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
von: Xiang, Maoyang, et al.
Veröffentlicht: (2026)
von: Xiang, Maoyang, et al.
Veröffentlicht: (2026)
Deep Orthogonal Hypersphere Compression for Anomaly Detection
von: Zhang, Yunhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yunhe, et al.
Veröffentlicht: (2023)
Integration of Mamba and Transformer -- MAT for Long-Short Range Time Series Forecasting with Application to Weather Dynamics
von: Zhang, Wenqing, et al.
Veröffentlicht: (2024)
von: Zhang, Wenqing, et al.
Veröffentlicht: (2024)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
Multiplicative Orthogonal Sequential Editing for Language Models
von: Xu, Hao-Xiang, et al.
Veröffentlicht: (2026)
von: Xu, Hao-Xiang, et al.
Veröffentlicht: (2026)
Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality Regularization
von: He, Junlin, et al.
Veröffentlicht: (2024)
von: He, Junlin, et al.
Veröffentlicht: (2024)
Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
von: Yang, Hanlin, et al.
Veröffentlicht: (2024)
von: Yang, Hanlin, et al.
Veröffentlicht: (2024)
ROOT: Robust Orthogonalized Optimizer for Neural Network Training
von: He, Wei, et al.
Veröffentlicht: (2025)
von: He, Wei, et al.
Veröffentlicht: (2025)
Recovering implicit physics model under real-world constraints
von: Banerjee, Ayan, et al.
Veröffentlicht: (2024)
von: Banerjee, Ayan, et al.
Veröffentlicht: (2024)
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning
von: Zhang, Yixian, et al.
Veröffentlicht: (2025)
von: Zhang, Yixian, et al.
Veröffentlicht: (2025)
Pre-training with Random Orthogonal Projection Image Modeling
von: Haghighat, Maryam, et al.
Veröffentlicht: (2023)
von: Haghighat, Maryam, et al.
Veröffentlicht: (2023)
Scalable Heterogeneous Graph Learning via Heterogeneous-aware Orthogonal Prototype Experts
von: Zhou, Wei, et al.
Veröffentlicht: (2026)
von: Zhou, Wei, et al.
Veröffentlicht: (2026)
Disentangling Task Conflicts in Multi-Task LoRA via Orthogonal Gradient Projection
von: Yang, Ziyu, et al.
Veröffentlicht: (2026)
von: Yang, Ziyu, et al.
Veröffentlicht: (2026)
KL for a KL: On-Policy Distillation with Control Variate Baseline
von: Oh, Minjae, et al.
Veröffentlicht: (2026)
von: Oh, Minjae, et al.
Veröffentlicht: (2026)
Recovering from Biased Data: Can Fairness Constraints Improve Accuracy?
von: Blum, Avrim, et al.
Veröffentlicht: (2019)
von: Blum, Avrim, et al.
Veröffentlicht: (2019)
Towards the Transferability of Rewards Recovered via Regularized Inverse Reinforcement Learning
von: Schlaginhaufen, Andreas, et al.
Veröffentlicht: (2024)
von: Schlaginhaufen, Andreas, et al.
Veröffentlicht: (2024)
Optimal Stability of KL Divergence under Gaussian Perturbations
von: Pan, Jialu, et al.
Veröffentlicht: (2026)
von: Pan, Jialu, et al.
Veröffentlicht: (2026)
KL-regularization Itself is Differentially Private in Bandits and RLHF
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
ORTHOBO: Orthogonal Bayesian Hyperparameter Optimization
von: Schröder, Maresa, et al.
Veröffentlicht: (2026)
von: Schröder, Maresa, et al.
Veröffentlicht: (2026)
Measuring Orthogonality in Representations of Generative Models
von: Geyer, Robin C., et al.
Veröffentlicht: (2024)
von: Geyer, Robin C., et al.
Veröffentlicht: (2024)
ONG: Orthogonal Natural Gradient Descent
von: Yadav, Yajat, et al.
Veröffentlicht: (2025)
von: Yadav, Yajat, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2025) -
In-Context Reward Adaptation for Robust Preference Modeling
von: Sun, Zhenyu, et al.
Veröffentlicht: (2026) -
SOAP: Improving and Stabilizing Shampoo using Adam
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024) -
4-bit Shampoo for Memory-Efficient Network Training
von: Wang, Sike, et al.
Veröffentlicht: (2024) -
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2026)