Memory-Efficient LLM Training with Online Subspace Descent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Kaizhao, Liu, Bo, Chen, Lizhang, Liu, Qiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cautious Optimizers: Improving Training with One Line of Code
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts
von: Chen, Lizhang, et al.
Veröffentlicht: (2023)
von: Chen, Lizhang, et al.
Veröffentlicht: (2023)
Communication Efficient Distributed Training with Distributed Lion
von: Liu, Bo, et al.
Veröffentlicht: (2024)
von: Liu, Bo, et al.
Veröffentlicht: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
Memory-Efficient Optimization with Factorized Hamiltonian Descent
von: Nguyen, Son, et al.
Veröffentlicht: (2024)
von: Nguyen, Son, et al.
Veröffentlicht: (2024)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
SABER: Switchable and Balanced Training for Efficient LLM Reasoning
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
DRESSing Up LLM: Efficient Stylized Question-Answering via Style Subspace Editing
von: Ma, Xinyu, et al.
Veröffentlicht: (2025)
von: Ma, Xinyu, et al.
Veröffentlicht: (2025)
MeKi: Memory-based Expert Knowledge Injection for Efficient LLM Scaling
von: Ding, Ning, et al.
Veröffentlicht: (2026)
von: Ding, Ning, et al.
Veröffentlicht: (2026)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
COAP: Memory-Efficient Training with Correlation-Aware Gradient Projection
von: Xiao, Jinqi, et al.
Veröffentlicht: (2024)
von: Xiao, Jinqi, et al.
Veröffentlicht: (2024)
DB-LLM: Accurate Dual-Binarization for Efficient LLMs
von: Chen, Hong, et al.
Veröffentlicht: (2024)
von: Chen, Hong, et al.
Veröffentlicht: (2024)
AdaFRUGAL: Adaptive Memory-Efficient Training with Dynamic Control
von: Bui, Quang-Hung, et al.
Veröffentlicht: (2025)
von: Bui, Quang-Hung, et al.
Veröffentlicht: (2025)
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
von: Nair, Pranav Ajit, et al.
Veröffentlicht: (2024)
von: Nair, Pranav Ajit, et al.
Veröffentlicht: (2024)
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
von: Hameed, Marawan Gamal Abdel, et al.
Veröffentlicht: (2024)
von: Hameed, Marawan Gamal Abdel, et al.
Veröffentlicht: (2024)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
von: Wu, Yongtao, et al.
Veröffentlicht: (2025)
von: Wu, Yongtao, et al.
Veröffentlicht: (2025)
Muon is Scalable for LLM Training
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
ArcMemo: Abstract Reasoning Composition with Lifelong LLM Memory
von: Ho, Matthew, et al.
Veröffentlicht: (2025)
von: Ho, Matthew, et al.
Veröffentlicht: (2025)
Scaling with Collapse: Efficient and Predictable Training of LLM Families
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
Efficient Agent Training for Computer Use
von: He, Yanheng, et al.
Veröffentlicht: (2025)
von: He, Yanheng, et al.
Veröffentlicht: (2025)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2023)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2023)
SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
von: Huang, Tianjin, et al.
Veröffentlicht: (2025)
von: Huang, Tianjin, et al.
Veröffentlicht: (2025)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Composition of Experts: A Modular Compound AI System Leveraging Large Language Models
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
Joint Detection of Fraud and Concept Drift inOnline Conversations with LLM-Assisted Judgment
von: Senol, Ali, et al.
Veröffentlicht: (2025)
von: Senol, Ali, et al.
Veröffentlicht: (2025)
CoRA: Optimizing Low-Rank Adaptation with Common Subspace of Large Language Models
von: Xiao, Xiaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Xiaojun, et al.
Veröffentlicht: (2024)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
von: Zhao, Haiyan, et al.
Veröffentlicht: (2024)
von: Zhao, Haiyan, et al.
Veröffentlicht: (2024)
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
von: Xu, Chejian, et al.
Veröffentlicht: (2025)
von: Xu, Chejian, et al.
Veröffentlicht: (2025)
User-LLM: Efficient LLM Contextualization with User Embeddings
von: Ning, Lin, et al.
Veröffentlicht: (2024)
von: Ning, Lin, et al.
Veröffentlicht: (2024)
Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
von: Dade, Nii Osae Osae, et al.
Veröffentlicht: (2025)
von: Dade, Nii Osae Osae, et al.
Veröffentlicht: (2025)
RecMem: Recurrence-based Memory Consolidation for Efficient and Effective Long-Running LLM Agents
von: Dai, Zijie, et al.
Veröffentlicht: (2026)
von: Dai, Zijie, et al.
Veröffentlicht: (2026)
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
von: Yu, Hongli, et al.
Veröffentlicht: (2025)
von: Yu, Hongli, et al.
Veröffentlicht: (2025)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2025)
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2025)
Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
von: Gan, Zeyu, et al.
Veröffentlicht: (2024)
von: Gan, Zeyu, et al.
Veröffentlicht: (2024)
Understanding In-context Learning of Addition via Activation Subspaces
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
von: Zhou, Cai, et al.
Veröffentlicht: (2026)
von: Zhou, Cai, et al.
Veröffentlicht: (2026)
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
von: Lai, Kunfeng, et al.
Veröffentlicht: (2025)
von: Lai, Kunfeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Cautious Optimizers: Improving Training with One Line of Code
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024) -
Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts
von: Chen, Lizhang, et al.
Veröffentlicht: (2023) -
Communication Efficient Distributed Training with Distributed Lion
von: Liu, Bo, et al.
Veröffentlicht: (2024) -
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024) -
Memory-Efficient Optimization with Factorized Hamiltonian Descent
von: Nguyen, Son, et al.
Veröffentlicht: (2024)