SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Refael, Yehonathan, Smorodinsky, Guy, Tirer, Tom, Lindenbaum, Ofir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LORENZA: Enhancing Generalization in Low-Rank Gradient LLM Training via Efficient Zeroth-Order Adaptive SAM
von: Refael, Yehonathan, et al.
Veröffentlicht: (2025)
von: Refael, Yehonathan, et al.
Veröffentlicht: (2025)
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
von: Arbel, Iftach, et al.
Veröffentlicht: (2024)
von: Arbel, Iftach, et al.
Veröffentlicht: (2024)
Unveiling Multiple Descents in Unsupervised Autoencoders
von: Rahimi, Kobi, et al.
Veröffentlicht: (2024)
von: Rahimi, Kobi, et al.
Veröffentlicht: (2024)
Train Less, Infer Faster: Efficient Model Finetuning and Compression via Structured Sparsity
von: Svirsky, Jonathan, et al.
Veröffentlicht: (2026)
von: Svirsky, Jonathan, et al.
Veröffentlicht: (2026)
AdaRankGrad: Adaptive Gradient-Rank and Moments for Memory-Efficient LLMs Training and Fine-Tuning
von: Refael, Yehonathan, et al.
Veröffentlicht: (2024)
von: Refael, Yehonathan, et al.
Veröffentlicht: (2024)
No Prior, No Leakage: Revisiting Reconstruction Attacks in Trained Neural Networks
von: Refael, Yehonatan, et al.
Veröffentlicht: (2025)
von: Refael, Yehonatan, et al.
Veröffentlicht: (2025)
FineGates: LLMs Finetuning with Compression using Stochastic Gates
von: Svirsky, Jonathan, et al.
Veröffentlicht: (2024)
von: Svirsky, Jonathan, et al.
Veröffentlicht: (2024)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
von: Pan, Rui, et al.
Veröffentlicht: (2024)
von: Pan, Rui, et al.
Veröffentlicht: (2024)
Absorber LLM: Harnessing Causal Synchronization for Test-Time Training
von: Zhang, Zhixin, et al.
Veröffentlicht: (2026)
von: Zhang, Zhixin, et al.
Veröffentlicht: (2026)
Secure LLM Fine-Tuning via Safety-Aware Probing
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
Stochastic Subspace Descent Accelerated via Bi-fidelity Line Search
von: Cheng, Nuojin, et al.
Veröffentlicht: (2025)
von: Cheng, Nuojin, et al.
Veröffentlicht: (2025)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
von: Zhang, Zhixin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhixin, et al.
Veröffentlicht: (2025)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
von: Yang, Tong, et al.
Veröffentlicht: (2024)
von: Yang, Tong, et al.
Veröffentlicht: (2024)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
von: Lu, Yao, et al.
Veröffentlicht: (2026)
von: Lu, Yao, et al.
Veröffentlicht: (2026)
A Unified Understanding of Offline Data Selection and Online Self-refining Generation for Post-training LLMs
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
Distributional Surgery for Language Model Activations
von: Nguyen, Bao, et al.
Veröffentlicht: (2025)
von: Nguyen, Bao, et al.
Veröffentlicht: (2025)
FOCUS: First Order Concentrated Updating Scheme
von: Liu, Yizhou, et al.
Veröffentlicht: (2025)
von: Liu, Yizhou, et al.
Veröffentlicht: (2025)
Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2024)
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2024)
COS-DPO: Conditioned One-Shot Multi-Objective Fine-Tuning Framework
von: Ren, Yinuo, et al.
Veröffentlicht: (2024)
von: Ren, Yinuo, et al.
Veröffentlicht: (2024)
Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
von: Bu, Zhiqi, et al.
Veröffentlicht: (2026)
von: Bu, Zhiqi, et al.
Veröffentlicht: (2026)
Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
von: Kunstner, Frederik, et al.
Veröffentlicht: (2024)
von: Kunstner, Frederik, et al.
Veröffentlicht: (2024)
Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
von: Hanqing, Liu, et al.
Veröffentlicht: (2026)
von: Hanqing, Liu, et al.
Veröffentlicht: (2026)
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback
von: Wang, Yikai, et al.
Veröffentlicht: (2026)
von: Wang, Yikai, et al.
Veröffentlicht: (2026)
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
von: Liu, Hong, et al.
Veröffentlicht: (2023)
von: Liu, Hong, et al.
Veröffentlicht: (2023)
Causal LLM Routing: End-to-End Regret Minimization from Observational Data
von: Tsiourvas, Asterios, et al.
Veröffentlicht: (2025)
von: Tsiourvas, Asterios, et al.
Veröffentlicht: (2025)
Subspace Optimization for Efficient Federated Learning under Heterogeneous Data
von: Zhu, Shuchen, et al.
Veröffentlicht: (2026)
von: Zhu, Shuchen, et al.
Veröffentlicht: (2026)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
ControlAgent: Automating Control System Design via Novel Integration of LLM Agents and Domain Expertise
von: Guo, Xingang, et al.
Veröffentlicht: (2024)
von: Guo, Xingang, et al.
Veröffentlicht: (2024)
Efficient Optimization with Orthogonality Constraint: a Randomized Riemannian Submanifold Method
von: Han, Andi, et al.
Veröffentlicht: (2025)
von: Han, Andi, et al.
Veröffentlicht: (2025)
Faster Stochastic Optimization with Arbitrary Delays via Asynchronous Mini-Batching
von: Attia, Amit, et al.
Veröffentlicht: (2024)
von: Attia, Amit, et al.
Veröffentlicht: (2024)
Masked Subspace Clustering Methods
von: Song, Jiebo, et al.
Veröffentlicht: (2025)
von: Song, Jiebo, et al.
Veröffentlicht: (2025)
Memory-Efficient LLM Training with Online Subspace Descent
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
Nesterov Acceleration for Ensemble Kalman Inversion and Variants
von: Vernon, Sydney, et al.
Veröffentlicht: (2025)
von: Vernon, Sydney, et al.
Veröffentlicht: (2025)
Accelerated stochastic approximation with state-dependent noise
von: Ilandarideva, Sasila, et al.
Veröffentlicht: (2023)
von: Ilandarideva, Sasila, et al.
Veröffentlicht: (2023)
RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning
von: Wei, Zeming, et al.
Veröffentlicht: (2026)
von: Wei, Zeming, et al.
Veröffentlicht: (2026)
Inference of Online Newton Methods with Nesterov's Accelerated Sketching
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
Muon in Associative Memory Learning: Training Dynamics and Scaling Laws
von: Li, Binghui, et al.
Veröffentlicht: (2026)
von: Li, Binghui, et al.
Veröffentlicht: (2026)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
von: Glentis, Athanasios, et al.
Veröffentlicht: (2025)
von: Glentis, Athanasios, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LORENZA: Enhancing Generalization in Low-Rank Gradient LLM Training via Efficient Zeroth-Order Adaptive SAM
von: Refael, Yehonathan, et al.
Veröffentlicht: (2025) -
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
von: Arbel, Iftach, et al.
Veröffentlicht: (2024) -
Unveiling Multiple Descents in Unsupervised Autoencoders
von: Rahimi, Kobi, et al.
Veröffentlicht: (2024) -
Train Less, Infer Faster: Efficient Model Finetuning and Compression via Structured Sparsity
von: Svirsky, Jonathan, et al.
Veröffentlicht: (2026) -
AdaRankGrad: Adaptive Gradient-Rank and Moments for Memory-Efficient LLMs Training and Fine-Tuning
von: Refael, Yehonathan, et al.
Veröffentlicht: (2024)