GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sattarifard, Amirmohsen, Lavasani, Sepehr, Zhang, Kunlin, Rajabpour, Amirhossein, Xu, Hanlin, Sun, Fengyu, Hassanpour, Negar, Gao, Chao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910213461770240
author Sattarifard, Amirmohsen
Lavasani, Sepehr
Zhang, Kunlin
Rajabpour, Amirhossein
Xu, Hanlin
Sun, Fengyu
Hassanpour, Negar
Gao, Chao
author_facet Sattarifard, Amirmohsen
Lavasani, Sepehr
Zhang, Kunlin
Rajabpour, Amirhossein
Xu, Hanlin
Sun, Fengyu
Hassanpour, Negar
Gao, Chao
contents Inference-time sparsification is a promising path to deploy large language models (LLMs) on resource-constrained devices, yet existing training-free methods typically estimate feedforward network (FFN) neuron importance from the input prompt alone. We show this prompt-only signal is often unreliable, especially for short prompts and long-form decoding, leading to inaccurate masks and degraded generation fidelity. We propose GLASS, a plug-and-play, training-free framework that stabilizes dynamic FFN pruning by aggregating two complementary views of neuron criticality: local prompt-specific activations and a global model-intrinsic prior. GLASS fuses global and local signals via rank aggregation, yielding robust critical-neuron selection even when the prompt is short. We interpret GLASS as the maximum-a-posteriori consensus ranking under a permutation-based probabilistic model, providing a principled foundation for its weighted rank-aggregation rule. We apply GLASS to a diverse set of open-source LLMs, and show that it yields substantial improvements over prior training-free baselines in the challenging short-prompt, long-generation scenarios, achieving up to 45.10% lower perplexity and 25.73% lower KL divergence, while delivering significant on-device decoding speedup.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14302
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
Sattarifard, Amirmohsen
Lavasani, Sepehr
Zhang, Kunlin
Rajabpour, Amirhossein
Xu, Hanlin
Sun, Fengyu
Hassanpour, Negar
Gao, Chao
Machine Learning
Artificial Intelligence
Computation and Language
Inference-time sparsification is a promising path to deploy large language models (LLMs) on resource-constrained devices, yet existing training-free methods typically estimate feedforward network (FFN) neuron importance from the input prompt alone. We show this prompt-only signal is often unreliable, especially for short prompts and long-form decoding, leading to inaccurate masks and degraded generation fidelity. We propose GLASS, a plug-and-play, training-free framework that stabilizes dynamic FFN pruning by aggregating two complementary views of neuron criticality: local prompt-specific activations and a global model-intrinsic prior. GLASS fuses global and local signals via rank aggregation, yielding robust critical-neuron selection even when the prompt is short. We interpret GLASS as the maximum-a-posteriori consensus ranking under a permutation-based probabilistic model, providing a principled foundation for its weighted rank-aggregation rule. We apply GLASS to a diverse set of open-source LLMs, and show that it yields substantial improvements over prior training-free baselines in the challenging short-prompt, long-generation scenarios, achieving up to 45.10% lower perplexity and 25.73% lower KL divergence, while delivering significant on-device decoding speedup.
title GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.14302