Saved in:
| Main Authors: | Sun, Ruotong, Wei, Ermin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.06316 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
by: Eschenhagen, Runa, et al.
Published: (2025)
by: Eschenhagen, Runa, et al.
Published: (2025)
In-Context Reward Adaptation for Robust Preference Modeling
by: Sun, Zhenyu, et al.
Published: (2026)
by: Sun, Zhenyu, et al.
Published: (2026)
SOAP: Improving and Stabilizing Shampoo using Adam
by: Vyas, Nikhil, et al.
Published: (2024)
by: Vyas, Nikhil, et al.
Published: (2024)
4-bit Shampoo for Memory-Efficient Network Training
by: Wang, Sike, et al.
Published: (2024)
by: Wang, Sike, et al.
Published: (2024)
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
by: Eschenhagen, Runa, et al.
Published: (2026)
by: Eschenhagen, Runa, et al.
Published: (2026)
Data Whitening Improves Sparse Autoencoder Learning
by: Saraswatula, Ashwin, et al.
Published: (2025)
by: Saraswatula, Ashwin, et al.
Published: (2025)
Orthogonalized Policy Optimization:Policy Optimization as Orthogonal Projection in Hilbert Space
by: Zixian, Wang
Published: (2026)
by: Zixian, Wang
Published: (2026)
Whitening Not Recommended for Classification Tasks in LLMs
by: Forooghi, Ali, et al.
Published: (2024)
by: Forooghi, Ali, et al.
Published: (2024)
Hessian-Free Online Certified Unlearning
by: Qiao, Xinbao, et al.
Published: (2024)
by: Qiao, Xinbao, et al.
Published: (2024)
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
by: Ma, Chao, et al.
Published: (2024)
by: Ma, Chao, et al.
Published: (2024)
Group Orthogonalized Policy Optimization:Group Policy Optimization as Orthogonal Projection in Hilbert Space
by: Zixian, Wang
Published: (2026)
by: Zixian, Wang
Published: (2026)
Reservoir Subspace Injection for Online ICA under Top-n Whitening
by: Xiao, Wenjun, et al.
Published: (2026)
by: Xiao, Wenjun, et al.
Published: (2026)
FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo
by: Nam, Kyunghun, et al.
Published: (2026)
by: Nam, Kyunghun, et al.
Published: (2026)
Fisher-Orthogonal Projected Natural Gradient Descent for Continual Learning
by: Garg, Ishir, et al.
Published: (2026)
by: Garg, Ishir, et al.
Published: (2026)
Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
Orthogonal Subspace Projection for Continual Machine Unlearning via SVD-Based LoRA
by: Rahulamathavan, Yogachandran, et al.
Published: (2026)
by: Rahulamathavan, Yogachandran, et al.
Published: (2026)
Approximated Orthogonal Projection Unit: Stabilizing Regression Network Training Using Natural Gradient
by: Wang, Shaoqi, et al.
Published: (2024)
by: Wang, Shaoqi, et al.
Published: (2024)
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
by: Lu, Yishun, et al.
Published: (2025)
by: Lu, Yishun, et al.
Published: (2025)
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
by: Shu, Yao, et al.
Published: (2026)
by: Shu, Yao, et al.
Published: (2026)
Integration of Mamba and Transformer -- MAT for Long-Short Range Time Series Forecasting with Application to Weather Dynamics
by: Zhang, Wenqing, et al.
Published: (2024)
by: Zhang, Wenqing, et al.
Published: (2024)
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
by: Xiang, Maoyang, et al.
Published: (2026)
by: Xiang, Maoyang, et al.
Published: (2026)
Deep Orthogonal Hypersphere Compression for Anomaly Detection
by: Zhang, Yunhe, et al.
Published: (2023)
by: Zhang, Yunhe, et al.
Published: (2023)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
by: Lin, Bokai, et al.
Published: (2024)
by: Lin, Bokai, et al.
Published: (2024)
Multiplicative Orthogonal Sequential Editing for Language Models
by: Xu, Hao-Xiang, et al.
Published: (2026)
by: Xu, Hao-Xiang, et al.
Published: (2026)
Pre-training with Random Orthogonal Projection Image Modeling
by: Haghighat, Maryam, et al.
Published: (2023)
by: Haghighat, Maryam, et al.
Published: (2023)
Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality Regularization
by: He, Junlin, et al.
Published: (2024)
by: He, Junlin, et al.
Published: (2024)
Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
by: Yang, Hanlin, et al.
Published: (2024)
by: Yang, Hanlin, et al.
Published: (2024)
ROOT: Robust Orthogonalized Optimizer for Neural Network Training
by: He, Wei, et al.
Published: (2025)
by: He, Wei, et al.
Published: (2025)
Disentangling Task Conflicts in Multi-Task LoRA via Orthogonal Gradient Projection
by: Yang, Ziyu, et al.
Published: (2026)
by: Yang, Ziyu, et al.
Published: (2026)
Recovering implicit physics model under real-world constraints
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
Hebbian Learning based Orthogonal Projection for Continual Learning of Spiking Neural Networks
by: Xiao, Mingqing, et al.
Published: (2024)
by: Xiao, Mingqing, et al.
Published: (2024)
Scalable Heterogeneous Graph Learning via Heterogeneous-aware Orthogonal Prototype Experts
by: Zhou, Wei, et al.
Published: (2026)
by: Zhou, Wei, et al.
Published: (2026)
KL for a KL: On-Policy Distillation with Control Variate Baseline
by: Oh, Minjae, et al.
Published: (2026)
by: Oh, Minjae, et al.
Published: (2026)
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning
by: Zhang, Yixian, et al.
Published: (2025)
by: Zhang, Yixian, et al.
Published: (2025)
Recovering from Biased Data: Can Fairness Constraints Improve Accuracy?
by: Blum, Avrim, et al.
Published: (2019)
by: Blum, Avrim, et al.
Published: (2019)
Towards the Transferability of Rewards Recovered via Regularized Inverse Reinforcement Learning
by: Schlaginhaufen, Andreas, et al.
Published: (2024)
by: Schlaginhaufen, Andreas, et al.
Published: (2024)
Privacy-Aware Video Anomaly Detection through Orthogonal Subspace Projection
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
ORTHOBO: Orthogonal Bayesian Hyperparameter Optimization
by: Schröder, Maresa, et al.
Published: (2026)
by: Schröder, Maresa, et al.
Published: (2026)
Measuring Orthogonality in Representations of Generative Models
by: Geyer, Robin C., et al.
Published: (2024)
by: Geyer, Robin C., et al.
Published: (2024)
ONG: Orthogonal Natural Gradient Descent
by: Yadav, Yajat, et al.
Published: (2025)
by: Yadav, Yajat, et al.
Published: (2025)
Similar Items
-
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
by: Eschenhagen, Runa, et al.
Published: (2025) -
In-Context Reward Adaptation for Robust Preference Modeling
by: Sun, Zhenyu, et al.
Published: (2026) -
SOAP: Improving and Stabilizing Shampoo using Adam
by: Vyas, Nikhil, et al.
Published: (2024) -
4-bit Shampoo for Memory-Efficient Network Training
by: Wang, Sike, et al.
Published: (2024) -
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
by: Eschenhagen, Runa, et al.
Published: (2026)