Indirect Attention: Turning Context Misalignment into a Feature

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bahaduri, Bissmella, Talaoubrid, Hicham, Feng, Fangchen, Ming, Zuheng, Mokraoui, Anissa
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912617666183168
author Bahaduri, Bissmella
Talaoubrid, Hicham
Feng, Fangchen
Ming, Zuheng
Mokraoui, Anissa
author_facet Bahaduri, Bissmella
Talaoubrid, Hicham
Feng, Fangchen
Ming, Zuheng
Mokraoui, Anissa
contents The attention mechanism has become a cornerstone of modern deep learning architectures, where keys and values are typically derived from the same underlying sequence or representation. This work explores a less conventional scenario, when keys and values originate from different sequences or modalities. Specifically, we first analyze the attention mechanism's behavior under noisy value features, establishing a critical noise threshold beyond which signal degradation becomes significant. Furthermore, we model context (key, value) misalignment as an effective form of structured noise within the value features, demonstrating that the noise induced by such misalignment can substantially exceed this critical threshold, thereby compromising standard attention's efficacy. Motivated by this, we introduce Indirect Attention, a modified attention mechanism that infers relevance indirectly in scenarios with misaligned context. We evaluate the performance of Indirect Attention across a range of synthetic tasks and real world applications, showcasing its superior ability to handle misalignment.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26015
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Indirect Attention: Turning Context Misalignment into a Feature
Bahaduri, Bissmella
Talaoubrid, Hicham
Feng, Fangchen
Ming, Zuheng
Mokraoui, Anissa
Machine Learning
Artificial Intelligence
The attention mechanism has become a cornerstone of modern deep learning architectures, where keys and values are typically derived from the same underlying sequence or representation. This work explores a less conventional scenario, when keys and values originate from different sequences or modalities. Specifically, we first analyze the attention mechanism's behavior under noisy value features, establishing a critical noise threshold beyond which signal degradation becomes significant. Furthermore, we model context (key, value) misalignment as an effective form of structured noise within the value features, demonstrating that the noise induced by such misalignment can substantially exceed this critical threshold, thereby compromising standard attention's efficacy. Motivated by this, we introduce Indirect Attention, a modified attention mechanism that infers relevance indirectly in scenarios with misaligned context. We evaluate the performance of Indirect Attention across a range of synthetic tasks and real world applications, showcasing its superior ability to handle misalignment.
title Indirect Attention: Turning Context Misalignment into a Feature
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.26015