Saved in:
Bibliographic Details
Main Authors: Lee, Nakyung, Kim, Yeongoon, Oh, Minhae, Kim, Suhwan, Koo, Jin Woo, Jo, Hyewon, Lee, Jungwoo
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.07324
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916942236876800
author Lee, Nakyung
Kim, Yeongoon
Oh, Minhae
Kim, Suhwan
Koo, Jin Woo
Jo, Hyewon
Lee, Jungwoo
author_facet Lee, Nakyung
Kim, Yeongoon
Oh, Minhae
Kim, Suhwan
Koo, Jin Woo
Jo, Hyewon
Lee, Jungwoo
contents Transformer-based self-attention mechanism serves as the core of modern language models, yet it often suffers from localization, where attentions collapse onto a limited subset of tokens and fail to capture long-range dependencies. To address this issue, we propose Self-Attention One-step Belief Propagation (SAOBP), a refinement framework that injects multi-hop relationships through a belief propagation process. To interpret and quantify these interactions, we introduce Global Token Dependency (GTD) that captures the relative contribution of multihop connections within the attention graph. Empirical results indicate that SAOBP helps prevent entropy collapse in deeper layers and adaptively maintains GTD at task-appropriate levels, thereby supporting improvements in model performance. Importantly, we observe competitive gains in small-scale models, highlighting its potential for improving inference quality in resource-constrained scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07324
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
Lee, Nakyung
Kim, Yeongoon
Oh, Minhae
Kim, Suhwan
Koo, Jin Woo
Jo, Hyewon
Lee, Jungwoo
Computation and Language
Artificial Intelligence
Transformer-based self-attention mechanism serves as the core of modern language models, yet it often suffers from localization, where attentions collapse onto a limited subset of tokens and fail to capture long-range dependencies. To address this issue, we propose Self-Attention One-step Belief Propagation (SAOBP), a refinement framework that injects multi-hop relationships through a belief propagation process. To interpret and quantify these interactions, we introduce Global Token Dependency (GTD) that captures the relative contribution of multihop connections within the attention graph. Empirical results indicate that SAOBP helps prevent entropy collapse in deeper layers and adaptively maintains GTD at task-appropriate levels, thereby supporting improvements in model performance. Importantly, we observe competitive gains in small-scale models, highlighting its potential for improving inference quality in resource-constrained scenarios.
title Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.07324