Attention Sinks in Diffusion Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Rulli, Maximo Eduardo, Petruzzi, Simone, Michielon, Edoardo, Silvestri, Fabrizio, Scardapane, Simone, Devoto, Alessio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression
by: Devoto, Alessio, et al.
Published: (2024)
by: Devoto, Alessio, et al.
Published: (2024)
Adaptive Semantic Token Selection for AI-native Goal-oriented Communications
by: Devoto, Alessio, et al.
Published: (2024)
by: Devoto, Alessio, et al.
Published: (2024)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
by: Godey, Nathan, et al.
Published: (2025)
by: Godey, Nathan, et al.
Published: (2025)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
by: Devoto, Alessio, et al.
Published: (2025)
by: Devoto, Alessio, et al.
Published: (2025)
Efficient Streaming Language Models with Attention Sinks
by: Xiao, Guangxuan, et al.
Published: (2023)
by: Xiao, Guangxuan, et al.
Published: (2023)
Evaluating Latent Knowledge of Public Tabular Datasets in Large Language Models
by: Silvestri, Matteo, et al.
Published: (2025)
by: Silvestri, Matteo, et al.
Published: (2025)
Natural Language Counterfactual Explanations for Graphs Using Large Language Models
by: Giorgi, Flavio, et al.
Published: (2024)
by: Giorgi, Flavio, et al.
Published: (2024)
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
by: Sun, Shangwen, et al.
Published: (2026)
by: Sun, Shangwen, et al.
Published: (2026)
Sink-Aware Pruning for Diffusion Language Models
by: Myrzakhan, Aidar, et al.
Published: (2026)
by: Myrzakhan, Aidar, et al.
Published: (2026)
Adaptive Layer Selection for Efficient Vision Transformer Fine-Tuning
by: Devoto, Alessio, et al.
Published: (2024)
by: Devoto, Alessio, et al.
Published: (2024)
When Attention Sink Emerges in Language Models: An Empirical View
by: Gu, Xiangming, et al.
Published: (2024)
by: Gu, Xiangming, et al.
Published: (2024)
One Token Is Enough: Improving Diffusion Language Models with a Sink Token
by: Zhang, Zihou, et al.
Published: (2026)
by: Zhang, Zihou, et al.
Published: (2026)
Spectral Filters, Dark Signals, and Attention Sinks
by: Cancedda, Nicola
Published: (2024)
by: Cancedda, Nicola
Published: (2024)
On the Existence and Behavior of Secondary Attention Sinks
by: Wong, Jeffrey T. H., et al.
Published: (2025)
by: Wong, Jeffrey T. H., et al.
Published: (2025)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
by: Luo, Jiayun, et al.
Published: (2025)
by: Luo, Jiayun, et al.
Published: (2025)
Alice's Adventures in a Differentiable Wonderland -- Volume I, A Tour of the Land
by: Scardapane, Simone
Published: (2024)
by: Scardapane, Simone
Published: (2024)
Conditional computation in neural networks: principles and research trends
by: Scardapane, Simone, et al.
Published: (2024)
by: Scardapane, Simone, et al.
Published: (2024)
Class incremental learning with probability dampening and cascaded gated classifier
by: Pomponi, Jary, et al.
Published: (2024)
by: Pomponi, Jary, et al.
Published: (2024)
Large Language Models to Diffusion Finetuning
by: Cetin, Edoardo, et al.
Published: (2025)
by: Cetin, Edoardo, et al.
Published: (2025)
Meronymic Ontology Extraction via Large Language Models
by: Zhang, Dekai, et al.
Published: (2025)
by: Zhang, Dekai, et al.
Published: (2025)
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
by: Verdini, Francesco, et al.
Published: (2024)
by: Verdini, Francesco, et al.
Published: (2024)
What are you sinking? A geometric approach on attention sink
by: Ruscio, Valeria, et al.
Published: (2025)
by: Ruscio, Valeria, et al.
Published: (2025)
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
by: Zhang, Stephen, et al.
Published: (2025)
by: Zhang, Stephen, et al.
Published: (2025)
LLMs in Interpreting Legal Documents
by: Corbo, Simone
Published: (2025)
by: Corbo, Simone
Published: (2025)
Do Large Language Models Understand Word Senses?
by: Meconi, Domenico, et al.
Published: (2025)
by: Meconi, Domenico, et al.
Published: (2025)
Select, Label, Evaluate: Active Testing in NLP
by: Purificato, Antonio, et al.
Published: (2026)
by: Purificato, Antonio, et al.
Published: (2026)
Loose LIPS Sink Ships: Asking Questions in Battleship with Language-Informed Program Sampling
by: Grand, Gabriel, et al.
Published: (2024)
by: Grand, Gabriel, et al.
Published: (2024)
BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations
by: Giovannini, Simone, et al.
Published: (2025)
by: Giovannini, Simone, et al.
Published: (2025)
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
by: Alvetreti, Federico, et al.
Published: (2025)
by: Alvetreti, Federico, et al.
Published: (2025)
Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs
by: Guo, Yanzhu, et al.
Published: (2024)
by: Guo, Yanzhu, et al.
Published: (2024)
PONTE: Personalized Orchestration for Natural Language Trustworthy Explanations
by: Vineis, Vittoria, et al.
Published: (2026)
by: Vineis, Vittoria, et al.
Published: (2026)
Reversible Diffusion Decoding for Diffusion Language Models
by: Wang, Xinyun, et al.
Published: (2026)
by: Wang, Xinyun, et al.
Published: (2026)
Multiple-Choice Question Generation Using Large Language Models: Methodology and Educator Insights
by: Biancini, Giorgio, et al.
Published: (2025)
by: Biancini, Giorgio, et al.
Published: (2025)
Attention-Aligned Reasoning for Large Language Models
by: Zhang, Hongxiang, et al.
Published: (2025)
by: Zhang, Hongxiang, et al.
Published: (2025)
Cross-Attention Watermarking of Large Language Models
by: Baldassini, Folco Bertini, et al.
Published: (2024)
by: Baldassini, Folco Bertini, et al.
Published: (2024)
Linguistic Profiling of a Neural Language Model
by: Miaschi, Alessio, et al.
Published: (2020)
by: Miaschi, Alessio, et al.
Published: (2020)
Are Large Language Models Moral Hypocrites? A Study Based on Moral Foundations
by: Nunes, José Luiz, et al.
Published: (2024)
by: Nunes, José Luiz, et al.
Published: (2024)
Non-Monotonic Attention-based Read/Write Policy Learning for Simultaneous Translation
by: Ahmed, Zeeshan, et al.
Published: (2025)
by: Ahmed, Zeeshan, et al.
Published: (2025)
MASS: MoErging through Adaptive Subspace Selection
by: Crisostomi, Donato, et al.
Published: (2025)
by: Crisostomi, Donato, et al.
Published: (2025)
Latent Multi-Head Attention for Small Language Models
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Similar Items
-
A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression
by: Devoto, Alessio, et al.
Published: (2024) -
Adaptive Semantic Token Selection for AI-native Goal-oriented Communications
by: Devoto, Alessio, et al.
Published: (2024) -
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
by: Godey, Nathan, et al.
Published: (2025) -
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
by: Devoto, Alessio, et al.
Published: (2025) -
Efficient Streaming Language Models with Attention Sinks
by: Xiao, Guangxuan, et al.
Published: (2023)