Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Shuaiyi, Zhang, Zhisong, Wang, Yan, Zhu, Lei, Ma, Dongyang, Deng, Chenlong, Deng, Yang, Lam, Wai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917522147639296
author Li, Shuaiyi
Zhang, Zhisong
Wang, Yan
Zhu, Lei
Ma, Dongyang
Deng, Chenlong
Deng, Yang
Lam, Wai
author_facet Li, Shuaiyi
Zhang, Zhisong
Wang, Yan
Zhu, Lei
Ma, Dongyang
Deng, Chenlong
Deng, Yang
Lam, Wai
contents Block attention, which processes the input as separate blocks that cannot attend to one another, offers significant potential to improve KV cache reuse in long-context scenarios such as Retrieval-Augmented Generation (RAG). However, its broader application is hindered by two key challenges: the difficulty of segmenting input text into meaningful, self-contained blocks, and the inefficiency of existing block fine-tuning methods that risk degrading performance. To address these, we first construct SemanticSeg, a large and diverse semantic segmentation dataset containing over 30k instances across 16 categories-including books, code, web text, and conversations with text lengths ranging from 2k to 32k. Using this dataset, we train a lightweight segmenter to automatically partition text into human-instinct-aligned blocks with controllable granularity. Second, we propose block distillation, a training framework that is more efficient than block fine-tuning, which uses a frozen full-attention teacher model to guide the block-attention student. This framework integrates three novel components: block sink tokens to mitigate information loss at block boundaries, block dropout to leverage training signals from all blocks, and token-level loss weighting to focus learning on block-attention-sensitive tokens. Experiments across multiple models and benchmarks demonstrate that our segmenter outperforms heuristic and statistical baselines, and block distillation achieves near-full-attention performance under block attention, establishing a practical and scalable pathway for deploying block attention.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15913
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
Li, Shuaiyi
Zhang, Zhisong
Wang, Yan
Zhu, Lei
Ma, Dongyang
Deng, Chenlong
Deng, Yang
Lam, Wai
Computation and Language
Artificial Intelligence
Block attention, which processes the input as separate blocks that cannot attend to one another, offers significant potential to improve KV cache reuse in long-context scenarios such as Retrieval-Augmented Generation (RAG). However, its broader application is hindered by two key challenges: the difficulty of segmenting input text into meaningful, self-contained blocks, and the inefficiency of existing block fine-tuning methods that risk degrading performance. To address these, we first construct SemanticSeg, a large and diverse semantic segmentation dataset containing over 30k instances across 16 categories-including books, code, web text, and conversations with text lengths ranging from 2k to 32k. Using this dataset, we train a lightweight segmenter to automatically partition text into human-instinct-aligned blocks with controllable granularity. Second, we propose block distillation, a training framework that is more efficient than block fine-tuning, which uses a frozen full-attention teacher model to guide the block-attention student. This framework integrates three novel components: block sink tokens to mitigate information loss at block boundaries, block dropout to leverage training signals from all blocks, and token-level loss weighting to focus learning on block-attention-sensitive tokens. Experiments across multiple models and benchmarks demonstrate that our segmenter outperforms heuristic and statistical baselines, and block distillation achieves near-full-attention performance under block attention, establishing a practical and scalable pathway for deploying block attention.
title Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.15913