Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Mingyu, Mei, Kai, Xu, Wujiang, Sun, Mingjie, Tang, Ruixiang, Du, Mengnan, Liu, Zirui, Zhang, Yongfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918027492065280
author Jin, Mingyu
Mei, Kai
Xu, Wujiang
Sun, Mingjie
Tang, Ruixiang
Du, Mengnan
Liu, Zirui
Zhang, Yongfeng
author_facet Jin, Mingyu
Mei, Kai
Xu, Wujiang
Sun, Mingjie
Tang, Ruixiang
Du, Mengnan
Liu, Zirui
Zhang, Yongfeng
contents Large language models (LLMs) have achieved remarkable success in contextual knowledge understanding. In this paper, we show that these concentrated massive values consistently emerge in specific regions of attention queries (Q) and keys (K) while not having such patterns in values (V) in various modern transformer-based LLMs (Q, K, and V mean the representations output by the query, key, and value layers respectively). Through extensive experiments, we further demonstrate that these massive values play a critical role in interpreting contextual knowledge (knowledge obtained from the current context window) rather than in retrieving parametric knowledge stored within the model's parameters. Our further investigation of quantization strategies reveals that ignoring these massive values leads to a pronounced drop in performance on tasks requiring rich contextual understanding, aligning with our analysis. Finally, we trace the emergence of concentrated massive values and find that such concentration is caused by Rotary Positional Encoding (RoPE), which has appeared since the first layers. These findings shed new light on how Q and K operate in LLMs and offer practical insights for model design and optimization. The Code is Available at https://github.com/MingyuJ666/Rope_with_LLM.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01563
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding
Jin, Mingyu
Mei, Kai
Xu, Wujiang
Sun, Mingjie
Tang, Ruixiang
Du, Mengnan
Liu, Zirui
Zhang, Yongfeng
Computation and Language
Large language models (LLMs) have achieved remarkable success in contextual knowledge understanding. In this paper, we show that these concentrated massive values consistently emerge in specific regions of attention queries (Q) and keys (K) while not having such patterns in values (V) in various modern transformer-based LLMs (Q, K, and V mean the representations output by the query, key, and value layers respectively). Through extensive experiments, we further demonstrate that these massive values play a critical role in interpreting contextual knowledge (knowledge obtained from the current context window) rather than in retrieving parametric knowledge stored within the model's parameters. Our further investigation of quantization strategies reveals that ignoring these massive values leads to a pronounced drop in performance on tasks requiring rich contextual understanding, aligning with our analysis. Finally, we trace the emergence of concentrated massive values and find that such concentration is caused by Rotary Positional Encoding (RoPE), which has appeared since the first layers. These findings shed new light on how Q and K operate in LLMs and offer practical insights for model design and optimization. The Code is Available at https://github.com/MingyuJ666/Rope_with_LLM.
title Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding
topic Computation and Language
url https://arxiv.org/abs/2502.01563