SVD Contextual Sparsity Predictors for Fast LLM Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Serbin, Georgii, Koshkin, Kirill, Sun, Zhongao, Bistrigova, Anastasiya, Korikov, C. C.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917342309515264
author Serbin, Georgii
Koshkin, Kirill
Sun, Zhongao
Bistrigova, Anastasiya
Korikov, C. C.
author_facet Serbin, Georgii
Koshkin, Kirill
Sun, Zhongao
Bistrigova, Anastasiya
Korikov, C. C.
contents Contextual sparsity is one of the approaches used to reduce computational complexity in the inference process of large language models (LLMs). Existing techniques for efficient LLM inference acceleration based on contextual sparsity with minimal accuracy degradation require training sparse pattern predictors. This paper presents a framework for accelerating inference of ReGLU-based feed-forward networks (FFNs) within LLMs. The proposed framework provides a fast, training-free method for building sparse pattern predictors using truncation-aware singular value decomposition (SVD) of the gate projection matrix, along with a threshold calibration algorithm, and inference executors supporting conditional computation on CUDA and CANN devices. Experiments on three sparse LLMs with an average activation sparsity level of 90% in the FFNs demonstrate up to a 1.8x reduction in end-to-end decoding time while maintaining less than 1% degradation in benchmark scores on tasks involving complex math and code generation. This work advances the deployment of LLMs on edge devices.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14110
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SVD Contextual Sparsity Predictors for Fast LLM Inference
Serbin, Georgii
Koshkin, Kirill
Sun, Zhongao
Bistrigova, Anastasiya
Korikov, C. C.
Machine Learning
Contextual sparsity is one of the approaches used to reduce computational complexity in the inference process of large language models (LLMs). Existing techniques for efficient LLM inference acceleration based on contextual sparsity with minimal accuracy degradation require training sparse pattern predictors. This paper presents a framework for accelerating inference of ReGLU-based feed-forward networks (FFNs) within LLMs. The proposed framework provides a fast, training-free method for building sparse pattern predictors using truncation-aware singular value decomposition (SVD) of the gate projection matrix, along with a threshold calibration algorithm, and inference executors supporting conditional computation on CUDA and CANN devices. Experiments on three sparse LLMs with an average activation sparsity level of 90% in the FFNs demonstrate up to a 1.8x reduction in end-to-end decoding time while maintaining less than 1% degradation in benchmark scores on tasks involving complex math and code generation. This work advances the deployment of LLMs on edge devices.
title SVD Contextual Sparsity Predictors for Fast LLM Inference
topic Machine Learning
url https://arxiv.org/abs/2603.14110