DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hanzhi, Fan, Heng, Sha, Kewei, Huang, Yan, Feng, Yunhe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908406300803072
author Zhang, Hanzhi
Fan, Heng
Sha, Kewei
Huang, Yan
Feng, Yunhe
author_facet Zhang, Hanzhi
Fan, Heng
Sha, Kewei
Huang, Yan
Feng, Yunhe
contents Long-context understanding is crucial for many NLP applications, yet transformers struggle with efficiency due to the quadratic complexity of self-attention. Sparse attention methods alleviate this cost but often impose static, predefined masks, failing to capture heterogeneous attention patterns. This results in suboptimal token interactions, limiting adaptability and retrieval accuracy in long-sequence tasks. This work introduces a dynamic sparse attention mechanism that assigns adaptive masks at the attention-map level, preserving heterogeneous patterns across layers and heads. Unlike existing approaches, our method eliminates the need for fine-tuning and predefined mask structures while maintaining computational efficiency. By learning context-aware attention structures, it achieves high alignment with full-attention models, ensuring minimal performance degradation while reducing memory and compute overhead. This approach provides a scalable alternative to full attention, enabling the practical deployment of large-scale Large Language Models (LLMs) without sacrificing retrieval performance. DAM is available at: https://github.com/HanzhiZhang-Ulrica/DAM.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11104
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
Zhang, Hanzhi
Fan, Heng
Sha, Kewei
Huang, Yan
Feng, Yunhe
Computation and Language
Artificial Intelligence
Long-context understanding is crucial for many NLP applications, yet transformers struggle with efficiency due to the quadratic complexity of self-attention. Sparse attention methods alleviate this cost but often impose static, predefined masks, failing to capture heterogeneous attention patterns. This results in suboptimal token interactions, limiting adaptability and retrieval accuracy in long-sequence tasks. This work introduces a dynamic sparse attention mechanism that assigns adaptive masks at the attention-map level, preserving heterogeneous patterns across layers and heads. Unlike existing approaches, our method eliminates the need for fine-tuning and predefined mask structures while maintaining computational efficiency. By learning context-aware attention structures, it achieves high alignment with full-attention models, ensuring minimal performance degradation while reducing memory and compute overhead. This approach provides a scalable alternative to full attention, enabling the practical deployment of large-scale Large Language Models (LLMs) without sacrificing retrieval performance. DAM is available at: https://github.com/HanzhiZhang-Ulrica/DAM.
title DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.11104