A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yiming, Zhang, Yuan, Liu, Yin, Yuan, Kun, Wen, Zaiwen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures
by: Chen, Yiming, et al.
Published: (2024)
by: Chen, Yiming, et al.
Published: (2024)
Subspace Optimization for Large Language Models with Convergence Guarantees
by: He, Yutong, et al.
Published: (2024)
by: He, Yutong, et al.
Published: (2024)
BROS: Bias-Corrected Randomized Subspaces for Memory-Efficient Single-Loop Bilevel Optimization
by: Zhang, Hengrui, et al.
Published: (2026)
by: Zhang, Hengrui, et al.
Published: (2026)
Randomized Gradient Subspaces for Efficient Large Language Model Training
by: Rajabi, Sahar, et al.
Published: (2025)
by: Rajabi, Sahar, et al.
Published: (2025)
OptProver: Bridging Olympiad and Optimization through Continual Training in Formal Theorem Proving
by: Li, Chenyi, et al.
Published: (2026)
by: Li, Chenyi, et al.
Published: (2026)
PRAC: Principal-Random Subspace for LLM Activation Compression and Memory-Efficient Training
by: Li, Yanyi, et al.
Published: (2026)
by: Li, Yanyi, et al.
Published: (2026)
GWT: Scalable Optimizer State Compression for Large Language Model Training
by: Wen, Ziqing, et al.
Published: (2025)
by: Wen, Ziqing, et al.
Published: (2025)
Memory-Efficient LLM Training with Online Subspace Descent
by: Liang, Kaizhao, et al.
Published: (2024)
by: Liang, Kaizhao, et al.
Published: (2024)
Accelerating Optimization via Differentiable Stopping Time
by: Xie, Zhonglin, et al.
Published: (2025)
by: Xie, Zhonglin, et al.
Published: (2025)
Accelerating LLM Pre-Training through Flat-Direction Dynamics Enhancement
by: Zhu, Shuchen, et al.
Published: (2026)
by: Zhu, Shuchen, et al.
Published: (2026)
An Efficient Subspace Algorithm for Federated Learning on Heterogeneous Data
by: Zhang, Jiaojiao, et al.
Published: (2025)
by: Zhang, Jiaojiao, et al.
Published: (2025)
Beyond Language: Format-Agnostic Reasoning Subspaces in Large Language Models
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language Models
by: Luo, Qijun, et al.
Published: (2024)
by: Luo, Qijun, et al.
Published: (2024)
OASIS: Online Activation Subspace Learning for Memory-Efficient Training
by: Choudhary, Sakshi, et al.
Published: (2026)
by: Choudhary, Sakshi, et al.
Published: (2026)
Constructing Industrial-Scale Optimization Modeling Benchmark
by: Li, Zhong, et al.
Published: (2026)
by: Li, Zhong, et al.
Published: (2026)
Orthogonal Subspace Clustering: Enhancing High-Dimensional Data Analysis through Adaptive Dimensionality Reduction and Efficient Clustering
by: Wen, Qing-Yuan, et al.
Published: (2026)
by: Wen, Qing-Yuan, et al.
Published: (2026)
Non-Asymptotic Global Convergence of PPO-Clip
by: Liu, Yin, et al.
Published: (2025)
by: Liu, Yin, et al.
Published: (2025)
COSMOS: A Hybrid Adaptive Optimizer for Memory-Efficient Training of LLMs
by: Liu, Liming, et al.
Published: (2025)
by: Liu, Liming, et al.
Published: (2025)
Efficient Resource-Constrained Training of Transformers via Subspace Optimization
by: Nguyen, Le-Trung, et al.
Published: (2025)
by: Nguyen, Le-Trung, et al.
Published: (2025)
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
CE-LoRA: Computation-Efficient LoRA Fine-Tuning for Language Models
by: Chen, Guanduo, et al.
Published: (2025)
by: Chen, Guanduo, et al.
Published: (2025)
MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
by: Liu, Yuxi, et al.
Published: (2025)
by: Liu, Yuxi, et al.
Published: (2025)
Certifying Language Model Robustness with Fuzzed Randomized Smoothing: An Efficient Defense Against Backdoor Attacks
by: He, Bowei, et al.
Published: (2025)
by: He, Bowei, et al.
Published: (2025)
Lotus: Efficient LLM Training by Randomized Low-Rank Gradient Projection with Adaptive Subspace Switching
by: Miao, Tianhao, et al.
Published: (2026)
by: Miao, Tianhao, et al.
Published: (2026)
OptMATH: A Scalable Bidirectional Data Synthesis Framework for Optimization Modeling
by: Lu, Hongliang, et al.
Published: (2025)
by: Lu, Hongliang, et al.
Published: (2025)
Lower Bounds and Accelerated Algorithms in Distributed Stochastic Optimization with Communication Compression
by: He, Yutong, et al.
Published: (2023)
by: He, Yutong, et al.
Published: (2023)
Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training
by: Zhou, Wenjie, et al.
Published: (2026)
by: Zhou, Wenjie, et al.
Published: (2026)
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
by: Li, Zhouyang, et al.
Published: (2025)
by: Li, Zhouyang, et al.
Published: (2025)
Powering Up Zeroth-Order Training via Subspace Gradient Orthogonalization
by: Lang, Yicheng, et al.
Published: (2026)
by: Lang, Yicheng, et al.
Published: (2026)
Learning Scalable Model Soup on a Single GPU: An Efficient Subspace Training Strategy
by: Li, Tao, et al.
Published: (2024)
by: Li, Tao, et al.
Published: (2024)
SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM Training
by: Refael, Yehonathan, et al.
Published: (2025)
by: Refael, Yehonathan, et al.
Published: (2025)
Gauss-Newton Temporal Difference Learning with Nonlinear Function Approximation
by: Ke, Zhifa, et al.
Published: (2023)
by: Ke, Zhifa, et al.
Published: (2023)
Perturbations in the Orthogonal Complement Subspace for Efficient Out-of-Distribution Detection
by: Huang, Zhexiao, et al.
Published: (2025)
by: Huang, Zhexiao, et al.
Published: (2025)
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
by: Kong, Boao, et al.
Published: (2026)
by: Kong, Boao, et al.
Published: (2026)
BackSlash: Rate Constrained Optimized Training of Large Language Models
by: Wu, Jun, et al.
Published: (2025)
by: Wu, Jun, et al.
Published: (2025)
CoRA: Optimizing Low-Rank Adaptation with Common Subspace of Large Language Models
by: Xiao, Xiaojun, et al.
Published: (2024)
by: Xiao, Xiaojun, et al.
Published: (2024)
CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
by: Kong, Boao, et al.
Published: (2025)
by: Kong, Boao, et al.
Published: (2025)
Simultaneous Computation and Memory Efficient Zeroth-Order Optimizer for Fine-Tuning Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Echo: A Large Language Model with Temporal Episodic Memory
by: Liu, WenTao, et al.
Published: (2025)
by: Liu, WenTao, et al.
Published: (2025)
Mixture-of-Channels: Exploiting Sparse FFNs for Efficient LLMs Pre-Training and Inference
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Similar Items
-
Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures
by: Chen, Yiming, et al.
Published: (2024) -
Subspace Optimization for Large Language Models with Convergence Guarantees
by: He, Yutong, et al.
Published: (2024) -
BROS: Bias-Corrected Randomized Subspaces for Memory-Efficient Single-Loop Bilevel Optimization
by: Zhang, Hengrui, et al.
Published: (2026) -
Randomized Gradient Subspaces for Efficient Large Language Model Training
by: Rajabi, Sahar, et al.
Published: (2025) -
OptProver: Bridging Olympiad and Optimization through Continual Training in Formal Theorem Proving
by: Li, Chenyi, et al.
Published: (2026)