A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Kaiyuan, Qiao, Linbo, Liu, Baihui, Jiang, Gongqingjian, Li, Shanshan, Li, Dongsheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908539083030528
author Tian, Kaiyuan
Qiao, Linbo
Liu, Baihui
Jiang, Gongqingjian
Li, Shanshan
Li, Dongsheng
author_facet Tian, Kaiyuan
Qiao, Linbo
Liu, Baihui
Jiang, Gongqingjian
Li, Shanshan
Li, Dongsheng
contents Scientific research faces high costs and inefficiencies with traditional methods, but the rise of deep learning and large language models (LLMs) offers innovative solutions. This survey reviews transformer-based LLM applications across scientific fields such as biology, medicine, chemistry, and meteorology, underscoring their role in advancing research. However, the continuous expansion of model size has led to significant memory demands, hindering further development and application of LLMs for science. This survey systematically reviews and categorizes memory-efficient pre-training techniques for large-scale transformers, including algorithm-level, system-level, and hardware-software co-optimization. Using AlphaFold 2 as an example, we demonstrate how tailored memory optimization methods can reduce storage needs while preserving prediction accuracy. By bridging model efficiency and scientific application needs, we hope to provide insights for scalable and cost-effective LLM training in AI for science.
format Preprint
id arxiv_https___arxiv_org_abs_2501_11847
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science
Tian, Kaiyuan
Qiao, Linbo
Liu, Baihui
Jiang, Gongqingjian
Li, Shanshan
Li, Dongsheng
Machine Learning
Artificial Intelligence
Scientific research faces high costs and inefficiencies with traditional methods, but the rise of deep learning and large language models (LLMs) offers innovative solutions. This survey reviews transformer-based LLM applications across scientific fields such as biology, medicine, chemistry, and meteorology, underscoring their role in advancing research. However, the continuous expansion of model size has led to significant memory demands, hindering further development and application of LLMs for science. This survey systematically reviews and categorizes memory-efficient pre-training techniques for large-scale transformers, including algorithm-level, system-level, and hardware-software co-optimization. Using AlphaFold 2 as an example, we demonstrate how tailored memory optimization methods can reduce storage needs while preserving prediction accuracy. By bridging model efficiency and scientific application needs, we hope to provide insights for scalable and cost-effective LLM training in AI for science.
title A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2501.11847