StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Qijun, Li, Mengqi, Zhao, Lei, Li, Xiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training Long-Context LLMs Efficiently via Chunk-wise Optimization
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs
by: Mao, Yuzhen, et al.
Published: (2026)
by: Mao, Yuzhen, et al.
Published: (2026)
Dynamic Spectral Backpropagation for Efficient Neural Network Training
by: Muthuraman, Mannmohan
Published: (2025)
by: Muthuraman, Mannmohan
Published: (2025)
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
by: Li, Mengqi, et al.
Published: (2025)
by: Li, Mengqi, et al.
Published: (2025)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
2BP: 2-Stage Backpropagation
by: Rae, Christopher, et al.
Published: (2024)
by: Rae, Christopher, et al.
Published: (2024)
FOAM: Blocked State Folding for Memory-Efficient LLM Training
by: Wen, Ziqing, et al.
Published: (2025)
by: Wen, Ziqing, et al.
Published: (2025)
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
by: Li, Zhouyang, et al.
Published: (2025)
by: Li, Zhouyang, et al.
Published: (2025)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
Advancing On-Device Neural Network Training with TinyPropv2: Dynamic, Sparse, and Efficient Backpropagation
by: Rüb, Marcus, et al.
Published: (2024)
by: Rüb, Marcus, et al.
Published: (2024)
SAL: Selective Adaptive Learning for Backpropagation-Free Training with Sparsification
by: Liu, Fanping, et al.
Published: (2026)
by: Liu, Fanping, et al.
Published: (2026)
Logarithmic Memory Networks (LMNs): Efficient Long-Range Sequence Modeling for Resource-Constrained Environments
by: Taha, Mohamed A.
Published: (2025)
by: Taha, Mohamed A.
Published: (2025)
4-bit Shampoo for Memory-Efficient Network Training
by: Wang, Sike, et al.
Published: (2024)
by: Wang, Sike, et al.
Published: (2024)
Near-Optimal Online Deployment and Routing for Streaming LLMs
by: Li, Shaoang, et al.
Published: (2025)
by: Li, Shaoang, et al.
Published: (2025)
SpanGNN: Towards Memory-Efficient Graph Neural Networks via Spanning Subgraph Training
by: Gu, Xizhi, et al.
Published: (2024)
by: Gu, Xizhi, et al.
Published: (2024)
Stochastic Layer-wise Learning: Scalable and Efficient Alternative to Backpropagation
by: Yin, Bojian, et al.
Published: (2025)
by: Yin, Bojian, et al.
Published: (2025)
Beyond Backpropagation: Exploring Innovative Algorithms for Energy-Efficient Deep Neural Network Training
by: Spyra, Przemysław
Published: (2025)
by: Spyra, Przemysław
Published: (2025)
AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
by: Han, Zhenyu, et al.
Published: (2025)
by: Han, Zhenyu, et al.
Published: (2025)
CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling
by: Zhao, Runsong, et al.
Published: (2026)
by: Zhao, Runsong, et al.
Published: (2026)
Practical Boolean Backpropagation
by: Golbert, Simon
Published: (2025)
by: Golbert, Simon
Published: (2025)
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
by: Li, Dacheng, et al.
Published: (2023)
by: Li, Dacheng, et al.
Published: (2023)
A Real-time Multimodal Transformer Neural Network-powered Wildfire Forecasting System
by: Chen, Qijun, et al.
Published: (2025)
by: Chen, Qijun, et al.
Published: (2025)
A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science
by: Tian, Kaiyuan, et al.
Published: (2025)
by: Tian, Kaiyuan, et al.
Published: (2025)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling
by: Zhang, Dehao, et al.
Published: (2025)
by: Zhang, Dehao, et al.
Published: (2025)
Trained Persistent Memory for Frozen Decoder-Only LLMs
by: Jeong, Hong
Published: (2026)
by: Jeong, Hong
Published: (2026)
Utility-Aware Data Pricing: Token-Level Quality and Empirical Training Gain for LLMs
by: Xu, Minghui, et al.
Published: (2026)
by: Xu, Minghui, et al.
Published: (2026)
HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
by: Luo, Cheng, et al.
Published: (2025)
by: Luo, Cheng, et al.
Published: (2025)
Long Input Sequence Network for Long Time Series Forecasting
by: Ma, Chao, et al.
Published: (2024)
by: Ma, Chao, et al.
Published: (2024)
FlashOptim: Optimizers for Memory-Efficient Training
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
Easy Samples Are All You Need: Self-Evolving LLMs via Data-Efficient Reinforcement Learning
by: Yu, Zhiyin, et al.
Published: (2026)
by: Yu, Zhiyin, et al.
Published: (2026)
FlashSampling: Fast and Memory-Efficient Exact Sampling
by: Ruiz, Tomas, et al.
Published: (2026)
by: Ruiz, Tomas, et al.
Published: (2026)
Neuroscience-Inspired Memory Replay for Continual Learning: A Comparative Study of Predictive Coding and Backpropagation-Based Strategies
by: Nalagatla, Goutham, et al.
Published: (2025)
by: Nalagatla, Goutham, et al.
Published: (2025)
Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
by: Xiao, Qiao, et al.
Published: (2026)
by: Xiao, Qiao, et al.
Published: (2026)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards
by: Li, Fanxing, et al.
Published: (2025)
by: Li, Fanxing, et al.
Published: (2025)
Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks
by: Pink, Mathis, et al.
Published: (2024)
by: Pink, Mathis, et al.
Published: (2024)
Backpropagation-Free Metropolis-Adjusted Langevin Algorithm
by: Cobb, Adam D., et al.
Published: (2025)
by: Cobb, Adam D., et al.
Published: (2025)
Similar Items
-
Training Long-Context LLMs Efficiently via Chunk-wise Optimization
by: Li, Wenhao, et al.
Published: (2025) -
IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs
by: Mao, Yuzhen, et al.
Published: (2026) -
Dynamic Spectral Backpropagation for Efficient Neural Network Training
by: Muthuraman, Mannmohan
Published: (2025) -
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
by: Li, Mengqi, et al.
Published: (2025) -
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)