A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Kaiyuan, Qiao, Linbo, Liu, Baihui, Jiang, Gongqingjian, Li, Shanshan, Li, Dongsheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
by: Liu, Baihui, et al.
Published: (2026)
by: Liu, Baihui, et al.
Published: (2026)
GRASS: Gradient-based Adaptive Layer-wise Importance Sampling for Memory-efficient Large Language Model Fine-tuning
by: Tian, Kaiyuan, et al.
Published: (2026)
by: Tian, Kaiyuan, et al.
Published: (2026)
ParaDySe: A Parallel-Strategy Switching Framework for Dynamic Sequence Lengths in Transformer
by: Ou, Zhixin, et al.
Published: (2025)
by: Ou, Zhixin, et al.
Published: (2025)
FOAM: Blocked State Folding for Memory-Efficient LLM Training
by: Wen, Ziqing, et al.
Published: (2025)
by: Wen, Ziqing, et al.
Published: (2025)
Explanation-Guided Adversarial Training for Robust and Interpretable Models
by: Chen, Chao, et al.
Published: (2026)
by: Chen, Chao, et al.
Published: (2026)
Acceleration for Deep Reinforcement Learning using Parallel and Distributed Computing: A Survey
by: Liu, Zhihong, et al.
Published: (2024)
by: Liu, Zhihong, et al.
Published: (2024)
4-bit Shampoo for Memory-Efficient Network Training
by: Wang, Sike, et al.
Published: (2024)
by: Wang, Sike, et al.
Published: (2024)
StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
by: Luo, Qijun, et al.
Published: (2025)
by: Luo, Qijun, et al.
Published: (2025)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
by: Liu, Zeyuan, et al.
Published: (2026)
by: Liu, Zeyuan, et al.
Published: (2026)
Survey and Taxonomy: The Role of Data-Centric AI in Transformer-Based Time Series Forecasting
by: Xu, Jingjing, et al.
Published: (2024)
by: Xu, Jingjing, et al.
Published: (2024)
CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling
by: Zhao, Runsong, et al.
Published: (2026)
by: Zhao, Runsong, et al.
Published: (2026)
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
by: Luo, Xufang, et al.
Published: (2025)
by: Luo, Xufang, et al.
Published: (2025)
SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond
by: Zhu, Xiangyang, et al.
Published: (2026)
by: Zhu, Xiangyang, et al.
Published: (2026)
Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
by: Xiao, Qiao, et al.
Published: (2026)
by: Xiao, Qiao, et al.
Published: (2026)
Data-Efficient Training by Evolved Sampling
by: Cheng, Ziheng, et al.
Published: (2025)
by: Cheng, Ziheng, et al.
Published: (2025)
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
by: Li, Zhouyang, et al.
Published: (2025)
by: Li, Zhouyang, et al.
Published: (2025)
The Role of Transformer Models in Advancing Blockchain Technology: A Systematic Survey
by: Liu, Tianxu, et al.
Published: (2024)
by: Liu, Tianxu, et al.
Published: (2024)
GWT: Scalable Optimizer State Compression for Large Language Model Training
by: Wen, Ziqing, et al.
Published: (2025)
by: Wen, Ziqing, et al.
Published: (2025)
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Machine Unlearning in Generative AI: A Survey
by: Liu, Zheyuan, et al.
Published: (2024)
by: Liu, Zheyuan, et al.
Published: (2024)
Generative Pre-Trained Transformer for Symbolic Regression Base In-Context Reinforcement Learning
by: Li, Yanjie, et al.
Published: (2024)
by: Li, Yanjie, et al.
Published: (2024)
ContiFormer: Continuous-Time Transformer for Irregular Time Series Modeling
by: Chen, Yuqi, et al.
Published: (2024)
by: Chen, Yuqi, et al.
Published: (2024)
Missingness-aware Data Imputation via AI-powered Bayesian Generative Modeling
by: Liu, Qiao
Published: (2026)
by: Liu, Qiao
Published: (2026)
FlashOptim: Optimizers for Memory-Efficient Training
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
Efficient Generative Model Training via Embedded Representation Warmup
by: Liu, Deyuan, et al.
Published: (2025)
by: Liu, Deyuan, et al.
Published: (2025)
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
A Survey on Time-Series Pre-Trained Models
by: Ma, Qianli, et al.
Published: (2023)
by: Ma, Qianli, et al.
Published: (2023)
Memory-Efficient LLM Training with Online Subspace Descent
by: Liang, Kaizhao, et al.
Published: (2024)
by: Liang, Kaizhao, et al.
Published: (2024)
Efficient Training of Large-Scale AI Models Through Federated Mixture-of-Experts: A System-Level Approach
by: Chen, Xiaobing, et al.
Published: (2025)
by: Chen, Xiaobing, et al.
Published: (2025)
Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
by: Li, Dawei, et al.
Published: (2025)
by: Li, Dawei, et al.
Published: (2025)
SpanGNN: Towards Memory-Efficient Graph Neural Networks via Spanning Subgraph Training
by: Gu, Xizhi, et al.
Published: (2024)
by: Gu, Xizhi, et al.
Published: (2024)
Detecting Multimedia Generated by Large AI Models: A Survey
by: Lin, Li, et al.
Published: (2024)
by: Lin, Li, et al.
Published: (2024)
Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized Models
by: Shi, Yubin, et al.
Published: (2024)
by: Shi, Yubin, et al.
Published: (2024)
XGrad: Boosting Gradient-Based Optimizers With Weight Prediction
by: Guan, Lei, et al.
Published: (2023)
by: Guan, Lei, et al.
Published: (2023)
Flow Matching Meets Biology and Life Science: A Survey
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
Feature-Function Curvature Analysis: A Geometric Framework for Explaining Differentiable Models
by: Najafi, Hamed, et al.
Published: (2025)
by: Najafi, Hamed, et al.
Published: (2025)
FedProphet: Memory-Efficient Federated Adversarial Training via Robust and Consistent Cascade Learning
by: Tang, Minxue, et al.
Published: (2024)
by: Tang, Minxue, et al.
Published: (2024)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
by: Qiu, Zeju, et al.
Published: (2026)
by: Qiu, Zeju, et al.
Published: (2026)
LOOKAT: Lookup-Optimized Key-Attention for Memory-Efficient Transformers
by: Karmore, Aryan
Published: (2026)
by: Karmore, Aryan
Published: (2026)
CoMERA: Computing- and Memory-Efficient Training via Rank-Adaptive Tensor Optimization
by: Yang, Zi, et al.
Published: (2024)
by: Yang, Zi, et al.
Published: (2024)
Similar Items
-
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
by: Liu, Baihui, et al.
Published: (2026) -
GRASS: Gradient-based Adaptive Layer-wise Importance Sampling for Memory-efficient Large Language Model Fine-tuning
by: Tian, Kaiyuan, et al.
Published: (2026) -
ParaDySe: A Parallel-Strategy Switching Framework for Dynamic Sequence Lengths in Transformer
by: Ou, Zhixin, et al.
Published: (2025) -
FOAM: Blocked State Folding for Memory-Efficient LLM Training
by: Wen, Ziqing, et al.
Published: (2025) -
Explanation-Guided Adversarial Training for Robust and Interpretable Models
by: Chen, Chao, et al.
Published: (2026)