Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Mingkuan, Gao, Yide, Hu, Wentao, Chen, Suquan, Huang, Tianchen, An, Zhenhua, Chang, Zetao, Sun, Xiayu, Min, Yuheng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911739731247104
author Zhao, Mingkuan
Gao, Yide
Hu, Wentao
Chen, Suquan
Huang, Tianchen
An, Zhenhua
Chang, Zetao
Sun, Xiayu
Min, Yuheng
author_facet Zhao, Mingkuan
Gao, Yide
Hu, Wentao
Chen, Suquan
Huang, Tianchen
An, Zhenhua
Chang, Zetao
Sun, Xiayu
Min, Yuheng
contents Large Language Models (LLMs) frequently exhibit "contextual disregard" when faced with input evidence that conflicts with their internal parametric memory, leading to persistent factual hallucinations. Existing mitigation strategies primarily rely on suppressing specific neuron activations or employing computationally expensive contrastive decoding mechanisms, which often result in increased perplexity or significantly elevated inference latency. To address these limitations, we propose Resonant Context Anchoring (RCA), a lightweight inference-time intervention method grounded in the perspective of residual stream signal dynamics. RCA aims to resolve the signal attenuation of external evidence during its propagation through deep networks. The core mechanism involves the orthogonal decoupling of routing logic and information magnitude within the self-attention module. By utilizing raw pre-softmax attention scores as an instantaneous metric of semantic alignment, we construct a dynamic gain field via non-linear rectification to selectively amplify the norms of value vectors corresponding to context tokens, without altering the attention probability distribution. This mechanism effectively elevates the signal-to-noise ratio (SNR) of input evidence within the residual stream mixture, thereby robustly anchoring the generation trajectory to the truthful context during inference. Extensive experiments on the Llama-3 model series demonstrate that RCA significantly improves contextual faithfulness across multiple factual consistency and strong knowledge-conflict tasks, effectively suppressing parametric hallucinations. Furthermore, results confirm that as a training-free and computationally negligible plug-and-play module, RCA achieves a Pareto improvement in faithfulness and fluency while maintaining the model's general language understanding capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01923
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
Zhao, Mingkuan
Gao, Yide
Hu, Wentao
Chen, Suquan
Huang, Tianchen
An, Zhenhua
Chang, Zetao
Sun, Xiayu
Min, Yuheng
Computation and Language
Machine Learning
I.2.7
Large Language Models (LLMs) frequently exhibit "contextual disregard" when faced with input evidence that conflicts with their internal parametric memory, leading to persistent factual hallucinations. Existing mitigation strategies primarily rely on suppressing specific neuron activations or employing computationally expensive contrastive decoding mechanisms, which often result in increased perplexity or significantly elevated inference latency. To address these limitations, we propose Resonant Context Anchoring (RCA), a lightweight inference-time intervention method grounded in the perspective of residual stream signal dynamics. RCA aims to resolve the signal attenuation of external evidence during its propagation through deep networks. The core mechanism involves the orthogonal decoupling of routing logic and information magnitude within the self-attention module. By utilizing raw pre-softmax attention scores as an instantaneous metric of semantic alignment, we construct a dynamic gain field via non-linear rectification to selectively amplify the norms of value vectors corresponding to context tokens, without altering the attention probability distribution. This mechanism effectively elevates the signal-to-noise ratio (SNR) of input evidence within the residual stream mixture, thereby robustly anchoring the generation trajectory to the truthful context during inference. Extensive experiments on the Llama-3 model series demonstrate that RCA significantly improves contextual faithfulness across multiple factual consistency and strong knowledge-conflict tasks, effectively suppressing parametric hallucinations. Furthermore, results confirm that as a training-free and computationally negligible plug-and-play module, RCA achieves a Pareto improvement in faithfulness and fluency while maintaining the model's general language understanding capabilities.
title Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
topic Computation and Language
Machine Learning
I.2.7
url https://arxiv.org/abs/2606.01923