DiffuMask: Diffusion Language Model for Token-level Prompt Pruning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zheng, Caleb, Singh, Jyotika, Tu, Fang, Sun, Weiyi, Bharadwaj, Sujeeth, Benajiba, Yassine, Ravi, Sujith, Shlizerman, Eli, Roth, Dan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917390019723264
author Zheng, Caleb
Singh, Jyotika
Tu, Fang
Sun, Weiyi
Bharadwaj, Sujeeth
Benajiba, Yassine
Ravi, Sujith
Shlizerman, Eli
Roth, Dan
author_facet Zheng, Caleb
Singh, Jyotika
Tu, Fang
Sun, Weiyi
Bharadwaj, Sujeeth
Benajiba, Yassine
Ravi, Sujith
Shlizerman, Eli
Roth, Dan
contents In-Context Learning and Chain-of-Thought prompting improve reasoning in large language models (LLMs). These typically come at the cost of longer, more expensive prompts that may contain redundant information. Prompt compression based on pruning offers a practical solution, yet existing methods rely on sequential token removal which is computationally intensive. We present DiffuMask, a diffusion-based framework integrating hierarchical shot-level and token-level pruning signals, that enables rapid and parallel prompt pruning via iterative mask prediction. DiffuMask substantially accelerates the compression process via masking multiple tokens in each denoising step. It offers tunable control over retained content, preserving essential reasoning context and achieving up to 80\% prompt length reduction. Meanwhile, it maintains or improves accuracy across in-domain, out-of-domain, and cross-model settings. Our results show that DiffuMask provides a generalizable and controllable framework for prompt compression, facilitating faster and more reliable in-context reasoning in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06627
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DiffuMask: Diffusion Language Model for Token-level Prompt Pruning
Zheng, Caleb
Singh, Jyotika
Tu, Fang
Sun, Weiyi
Bharadwaj, Sujeeth
Benajiba, Yassine
Ravi, Sujith
Shlizerman, Eli
Roth, Dan
Computation and Language
In-Context Learning and Chain-of-Thought prompting improve reasoning in large language models (LLMs). These typically come at the cost of longer, more expensive prompts that may contain redundant information. Prompt compression based on pruning offers a practical solution, yet existing methods rely on sequential token removal which is computationally intensive. We present DiffuMask, a diffusion-based framework integrating hierarchical shot-level and token-level pruning signals, that enables rapid and parallel prompt pruning via iterative mask prediction. DiffuMask substantially accelerates the compression process via masking multiple tokens in each denoising step. It offers tunable control over retained content, preserving essential reasoning context and achieving up to 80\% prompt length reduction. Meanwhile, it maintains or improves accuracy across in-domain, out-of-domain, and cross-model settings. Our results show that DiffuMask provides a generalizable and controllable framework for prompt compression, facilitating faster and more reliable in-context reasoning in LLMs.
title DiffuMask: Diffusion Language Model for Token-level Prompt Pruning
topic Computation and Language
url https://arxiv.org/abs/2604.06627