EcoSpa: Efficient Transformer Training with Coupled Sparsity

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xiao, Jinqi, Luo, Cheng, Huang, Lingyi, Yang, Cheng, Sui, Yang, Phan, Huy, Zang, Xiao, Ying, Yibiao, Tang, Zhexiang, Anandkumar, Anima, Yuan, Bo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918202346307584
author Xiao, Jinqi
Luo, Cheng
Huang, Lingyi
Yang, Cheng
Sui, Yang
Phan, Huy
Zang, Xiao
Ying, Yibiao
Tang, Zhexiang
Anandkumar, Anima
Yuan, Bo
author_facet Xiao, Jinqi
Luo, Cheng
Huang, Lingyi
Yang, Cheng
Sui, Yang
Phan, Huy
Zang, Xiao
Ying, Yibiao
Tang, Zhexiang
Anandkumar, Anima
Yuan, Bo
contents Transformers have become the backbone of modern AI, yet their high computational demands pose critical system challenges. While sparse training offers efficiency gains, existing methods fail to preserve critical structural relationships between weight matrices that interact multiplicatively in attention and feed-forward layers. This oversight leads to performance degradation at high sparsity levels. We introduce EcoSpa, an efficient structured sparse training method that jointly evaluates and sparsifies coupled weight matrix pairs, preserving their interaction patterns through aligned row/column removal. EcoSpa introduces a new granularity for calibrating structural component importance and performs coupled estimation and sparsification across both pre-training and fine-tuning scenarios. Evaluations demonstrate substantial improvements: EcoSpa enables efficient training of LLaMA-1B with 50\% memory reduction and 21\% faster training, achieves $2.2\times$ model compression on GPT-2-Medium with $2.4$ lower perplexity, and delivers $1.6\times$ inference speedup. The approach uses standard PyTorch operations, requiring no custom hardware or kernels, making efficient transformer training accessible on commodity hardware.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11641
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EcoSpa: Efficient Transformer Training with Coupled Sparsity
Xiao, Jinqi
Luo, Cheng
Huang, Lingyi
Yang, Cheng
Sui, Yang
Phan, Huy
Zang, Xiao
Ying, Yibiao
Tang, Zhexiang
Anandkumar, Anima
Yuan, Bo
Machine Learning
Artificial Intelligence
Performance
Transformers have become the backbone of modern AI, yet their high computational demands pose critical system challenges. While sparse training offers efficiency gains, existing methods fail to preserve critical structural relationships between weight matrices that interact multiplicatively in attention and feed-forward layers. This oversight leads to performance degradation at high sparsity levels. We introduce EcoSpa, an efficient structured sparse training method that jointly evaluates and sparsifies coupled weight matrix pairs, preserving their interaction patterns through aligned row/column removal. EcoSpa introduces a new granularity for calibrating structural component importance and performs coupled estimation and sparsification across both pre-training and fine-tuning scenarios. Evaluations demonstrate substantial improvements: EcoSpa enables efficient training of LLaMA-1B with 50\% memory reduction and 21\% faster training, achieves $2.2\times$ model compression on GPT-2-Medium with $2.4$ lower perplexity, and delivers $1.6\times$ inference speedup. The approach uses standard PyTorch operations, requiring no custom hardware or kernels, making efficient transformer training accessible on commodity hardware.
title EcoSpa: Efficient Transformer Training with Coupled Sparsity
topic Machine Learning
Artificial Intelligence
Performance
url https://arxiv.org/abs/2511.11641