Train Less, Infer Faster: Efficient Model Finetuning and Compression via Structured Sparsity

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Svirsky, Jonathan, Refael, Yehonathan, Lindenbaum, Ofir
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914318126153728
author Svirsky, Jonathan
Refael, Yehonathan
Lindenbaum, Ofir
author_facet Svirsky, Jonathan
Refael, Yehonathan
Lindenbaum, Ofir
contents Fully finetuning foundation language models (LMs) with billions of parameters is often impractical due to high computational costs, memory requirements, and the risk of overfitting. Although methods like low-rank adapters help address these challenges by adding small trainable modules to the frozen LM, they also increase memory usage and do not reduce inference latency. We uncover an intriguing phenomenon: sparsifying specific model rows and columns enables efficient task adaptation without requiring weight tuning. We propose a scheme for effective finetuning via sparsification using training stochastic gates, which requires minimal trainable parameters, reduces inference time, and removes 20--40\% of model parameters without significant accuracy loss. Empirical results show it outperforms recent finetuning baselines in efficiency and performance. Additionally, we provide theoretical guarantees for the convergence of this stochastic gating process, and show that our method admits a simpler and better-conditioned optimization landscape compared to LoRA. Our results highlight sparsity as a compelling mechanism for task-specific adaptation in LMs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09169
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Train Less, Infer Faster: Efficient Model Finetuning and Compression via Structured Sparsity
Svirsky, Jonathan
Refael, Yehonathan
Lindenbaum, Ofir
Machine Learning
Fully finetuning foundation language models (LMs) with billions of parameters is often impractical due to high computational costs, memory requirements, and the risk of overfitting. Although methods like low-rank adapters help address these challenges by adding small trainable modules to the frozen LM, they also increase memory usage and do not reduce inference latency. We uncover an intriguing phenomenon: sparsifying specific model rows and columns enables efficient task adaptation without requiring weight tuning. We propose a scheme for effective finetuning via sparsification using training stochastic gates, which requires minimal trainable parameters, reduces inference time, and removes 20--40\% of model parameters without significant accuracy loss. Empirical results show it outperforms recent finetuning baselines in efficiency and performance. Additionally, we provide theoretical guarantees for the convergence of this stochastic gating process, and show that our method admits a simpler and better-conditioned optimization landscape compared to LoRA. Our results highlight sparsity as a compelling mechanism for task-specific adaptation in LMs.
title Train Less, Infer Faster: Efficient Model Finetuning and Compression via Structured Sparsity
topic Machine Learning
url https://arxiv.org/abs/2602.09169