FOAM: Blocked State Folding for Memory-Efficient LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Ziqing, Wang, Jiahuan, Luo, Ping, Li, Dongsheng, Sun, Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GWT: Scalable Optimizer State Compression for Large Language Model Training
by: Wen, Ziqing, et al.
Published: (2025)
by: Wen, Ziqing, et al.
Published: (2025)
Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio
by: Wen, Ziqing, et al.
Published: (2026)
by: Wen, Ziqing, et al.
Published: (2026)
Unveiling High-Probability Generalization in Decentralized SGD
by: Wang, Jiahuan, et al.
Published: (2026)
by: Wang, Jiahuan, et al.
Published: (2026)
Stability and Generalization for Decentralized Markov SGD
by: Wang, Jiahuan, et al.
Published: (2026)
by: Wang, Jiahuan, et al.
Published: (2026)
Local Gradient Regulation Stabilizes Federated Learning under Client Heterogeneity
by: Luo, Ping, et al.
Published: (2026)
by: Luo, Ping, et al.
Published: (2026)
A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science
by: Tian, Kaiyuan, et al.
Published: (2025)
by: Tian, Kaiyuan, et al.
Published: (2025)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
by: Liu, Zeyuan, et al.
Published: (2026)
by: Liu, Zeyuan, et al.
Published: (2026)
FoldAct: Efficient and Stable Context Folding for Long-Horizon Search Agents
by: Shao, Jiaqi, et al.
Published: (2025)
by: Shao, Jiaqi, et al.
Published: (2025)
MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization
by: Wang, Ziqing, et al.
Published: (2026)
by: Wang, Ziqing, et al.
Published: (2026)
HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
by: Luo, Cheng, et al.
Published: (2025)
by: Luo, Cheng, et al.
Published: (2025)
StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
by: Luo, Qijun, et al.
Published: (2025)
by: Luo, Qijun, et al.
Published: (2025)
Bridge-IF: Learning Inverse Protein Folding with Markov Bridges
by: Zhu, Yiheng, et al.
Published: (2024)
by: Zhu, Yiheng, et al.
Published: (2024)
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
by: Li, Zhouyang, et al.
Published: (2025)
by: Li, Zhouyang, et al.
Published: (2025)
VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
by: Yu, Jiahuan, et al.
Published: (2025)
by: Yu, Jiahuan, et al.
Published: (2025)
4-bit Shampoo for Memory-Efficient Network Training
by: Wang, Sike, et al.
Published: (2024)
by: Wang, Sike, et al.
Published: (2024)
Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
by: Xiao, Qiao, et al.
Published: (2026)
by: Xiao, Qiao, et al.
Published: (2026)
Memory-Efficient LLM Training with Online Subspace Descent
by: Liang, Kaizhao, et al.
Published: (2024)
by: Liang, Kaizhao, et al.
Published: (2024)
Exploiting Block Coordinate Descent for Cost-Effective LLM Model Training
by: Liu, Zeyu, et al.
Published: (2025)
by: Liu, Zeyu, et al.
Published: (2025)
State Contamination in Memory-Augmented LLM Agents
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
An Efficient Training Algorithm for Models with Block-wise Sparsity
by: Zhu, Ding, et al.
Published: (2025)
by: Zhu, Ding, et al.
Published: (2025)
LNN-PINN: A Unified Physics-Only Training Framework with Liquid Residual Blocks
by: Tao, Ze, et al.
Published: (2025)
by: Tao, Ze, et al.
Published: (2025)
Efficient Generative Model Training via Embedded Representation Warmup
by: Liu, Deyuan, et al.
Published: (2025)
by: Liu, Deyuan, et al.
Published: (2025)
Score-based Generative Models with Adaptive Momentum
by: Wen, Ziqing, et al.
Published: (2024)
by: Wen, Ziqing, et al.
Published: (2024)
Federated Prediction-Powered Inference from Decentralized Data
by: Luo, Ping, et al.
Published: (2024)
by: Luo, Ping, et al.
Published: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Efficient State Space Model via Fast Tensor Convolution and Block Diagonalization
by: Liang, Tongyi, et al.
Published: (2024)
by: Liang, Tongyi, et al.
Published: (2024)
StateLinFormer: Stateful Training Enhancing Long-term Memory in Navigation
by: Chen, Zhiyuan, et al.
Published: (2026)
by: Chen, Zhiyuan, et al.
Published: (2026)
FlashOptim: Optimizers for Memory-Efficient Training
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
State-Action Inpainting Diffuser for Continuous Control with Delay
by: Han, Dongqi, et al.
Published: (2026)
by: Han, Dongqi, et al.
Published: (2026)
RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
by: Liu, Ningyuan, et al.
Published: (2025)
by: Liu, Ningyuan, et al.
Published: (2025)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
by: Yu, Jiahuan, et al.
Published: (2026)
by: Yu, Jiahuan, et al.
Published: (2026)
Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
by: Abillama, Pierre, et al.
Published: (2025)
by: Abillama, Pierre, et al.
Published: (2025)
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
by: Lu, Guanxi, et al.
Published: (2025)
by: Lu, Guanxi, et al.
Published: (2025)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
by: Yang, Hanmei, et al.
Published: (2024)
by: Yang, Hanmei, et al.
Published: (2024)
AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
by: Han, Zhenyu, et al.
Published: (2025)
by: Han, Zhenyu, et al.
Published: (2025)
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
by: Luo, Xufang, et al.
Published: (2025)
by: Luo, Xufang, et al.
Published: (2025)
MemReward: Graph-Based Experience Memory for LLM Reward Prediction with Limited Labels
by: Luo, Tianyang, et al.
Published: (2026)
by: Luo, Tianyang, et al.
Published: (2026)
FedProphet: Memory-Efficient Federated Adversarial Training via Robust and Consistent Cascade Learning
by: Tang, Minxue, et al.
Published: (2024)
by: Tang, Minxue, et al.
Published: (2024)
Controlled LLM Training on Spectral Sphere
by: Xie, Tian, et al.
Published: (2026)
by: Xie, Tian, et al.
Published: (2026)
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
by: Chen, Mengzhao, et al.
Published: (2024)
by: Chen, Mengzhao, et al.
Published: (2024)
Similar Items
-
GWT: Scalable Optimizer State Compression for Large Language Model Training
by: Wen, Ziqing, et al.
Published: (2025) -
Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio
by: Wen, Ziqing, et al.
Published: (2026) -
Unveiling High-Probability Generalization in Decentralized SGD
by: Wang, Jiahuan, et al.
Published: (2026) -
Stability and Generalization for Decentralized Markov SGD
by: Wang, Jiahuan, et al.
Published: (2026) -
Local Gradient Regulation Stabilizes Federated Learning under Client Heterogeneity
by: Luo, Ping, et al.
Published: (2026)