Measuring Affinity between Attention-Head Weight Subspaces via the Projection Kernel
Fuente:
arXiv
Saved in:
| Main Authors: | Yamagiwa, Hiroaki, Takase, Yusuke, Shimodaira, Hidetoshi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Axis Tour: Word Tour Determines the Order of Axes in ICA-transformed Embeddings
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
Mapping 1,000+ Language Models via the Log-Likelihood Vector
by: Oyama, Momose, et al.
Published: (2025)
by: Oyama, Momose, et al.
Published: (2025)
Norm of Mean Contextualized Embeddings Determines their Variance
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
Establishing a Scale for Kullback-Leibler Divergence in Language Models Across Various Settings
by: Kishino, Ryo, et al.
Published: (2025)
by: Kishino, Ryo, et al.
Published: (2025)
Revisiting Cosine Similarity via Normalized ICA-transformed Embeddings
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
Shimo Lab at "Discharge Me!": Discharge Summarization by Prompt-Driven Concatenation of Electronic Health Record Sections
by: He, Yunzhen, et al.
Published: (2024)
by: He, Yunzhen, et al.
Published: (2024)
Understanding Higher-Order Correlations Among Semantic Components in Embeddings
by: Oyama, Momose, et al.
Published: (2024)
by: Oyama, Momose, et al.
Published: (2024)
Language Model Maps for Prompt-Response Distributions via Log-Likelihood Vectors
by: Takase, Yusuke, et al.
Published: (2026)
by: Takase, Yusuke, et al.
Published: (2026)
Likelihood Variance as Text Importance for Resampling Texts to Map Language Models
by: Oyama, Momose, et al.
Published: (2025)
by: Oyama, Momose, et al.
Published: (2025)
Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport
by: Kishino, Ryo, et al.
Published: (2024)
by: Kishino, Ryo, et al.
Published: (2024)
Domain Mixture Design via Log-Likelihood Differences for Aligning Language Models with a Target Model
by: Kishino, Ryo, et al.
Published: (2026)
by: Kishino, Ryo, et al.
Published: (2026)
DeLTa: A Decoding Strategy based on Logit Trajectory Prediction Improves Factuality and Reasoning Ability
by: He, Yunzhen, et al.
Published: (2025)
by: He, Yunzhen, et al.
Published: (2025)
Predicting drug-gene relations via analogy tasks with word embeddings
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
Knowledge Sanitization of Large Language Models
by: Ishibashi, Yoichi, et al.
Published: (2023)
by: Ishibashi, Yoichi, et al.
Published: (2023)
3D Rotation and Translation for Hyperbolic Knowledge Graph Embedding
by: Zhu, Yihua, et al.
Published: (2023)
by: Zhu, Yihua, et al.
Published: (2023)
Block-Diagonal Orthogonal Relation and Matrix Entity for Knowledge Graph Embedding
by: Zhu, Yihua, et al.
Published: (2024)
by: Zhu, Yihua, et al.
Published: (2024)
Zipfian Whitening
by: Yokoi, Sho, et al.
Published: (2024)
by: Yokoi, Sho, et al.
Published: (2024)
Beyond Chains: Bridging Large Language Models and Knowledge Bases in Complex Question Answering
by: Zhu, Yihua, et al.
Published: (2025)
by: Zhu, Yihua, et al.
Published: (2025)
Self-Policy Distillation via Capability-Selective Subspace Projection
by: Hao, Guangya, et al.
Published: (2026)
by: Hao, Guangya, et al.
Published: (2026)
HARP: Hallucination Detection via Reasoning Subspace Projection
by: Hu, Junjie, et al.
Published: (2025)
by: Hu, Junjie, et al.
Published: (2025)
Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMs
by: Zhu, Yihua, et al.
Published: (2026)
by: Zhu, Yihua, et al.
Published: (2026)
Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks
by: Zhu, Yihua, et al.
Published: (2026)
by: Zhu, Yihua, et al.
Published: (2026)
Model Unlearning via Sparse Autoencoder Subspace Guided Projections
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Knocking-Heads Attention
by: Zhou, Zhanchao, et al.
Published: (2025)
by: Zhou, Zhanchao, et al.
Published: (2025)
Self-Translate-Train: Enhancing Cross-Lingual Transfer of Large Language Models via Inherent Capability
by: Ri, Ryokan, et al.
Published: (2024)
by: Ri, Ryokan, et al.
Published: (2024)
Toward LLMs Beyond English-Centric Development
by: Takase, Sho, et al.
Published: (2026)
by: Takase, Sho, et al.
Published: (2026)
AIR: Post-training Data Selection for Reasoning via Attention Head Influence
by: Liu, Jinrui, et al.
Published: (2025)
by: Liu, Jinrui, et al.
Published: (2025)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
by: Lin, Xihui, et al.
Published: (2024)
by: Lin, Xihui, et al.
Published: (2024)
ProxyAttn: Guided Sparse Attention via Representative Heads
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Contextual Subspace Manifold Projection for Structural Refinement of Large Language Model Representations
by: Wren, Alistair, et al.
Published: (2025)
by: Wren, Alistair, et al.
Published: (2025)
Interpreting Transformers Through Attention Head Intervention
by: Kadem, Mason, et al.
Published: (2026)
by: Kadem, Mason, et al.
Published: (2026)
Inferring Functionality of Attention Heads from their Parameters
by: Elhelo, Amit, et al.
Published: (2024)
by: Elhelo, Amit, et al.
Published: (2024)
Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Acceleration Multiple Heads Decoding for LLM via Dynamic Tree Attention
by: Zhang, Zhendong
Published: (2025)
by: Zhang, Zhendong
Published: (2025)
Head-wise Shareable Attention for Large Language Models
by: Cao, Zouying, et al.
Published: (2024)
by: Cao, Zouying, et al.
Published: (2024)
Attention Heads of Large Language Models: A Survey
by: Zheng, Zifan, et al.
Published: (2024)
by: Zheng, Zifan, et al.
Published: (2024)
Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads
by: He, Xingyang, et al.
Published: (2025)
by: He, Xingyang, et al.
Published: (2025)
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
by: Hua, Kai, et al.
Published: (2025)
by: Hua, Kai, et al.
Published: (2025)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
Similar Items
-
Axis Tour: Word Tour Determines the Order of Axes in ICA-transformed Embeddings
by: Yamagiwa, Hiroaki, et al.
Published: (2024) -
Mapping 1,000+ Language Models via the Log-Likelihood Vector
by: Oyama, Momose, et al.
Published: (2025) -
Norm of Mean Contextualized Embeddings Determines their Variance
by: Yamagiwa, Hiroaki, et al.
Published: (2024) -
Establishing a Scale for Kullback-Leibler Divergence in Language Models Across Various Settings
by: Kishino, Ryo, et al.
Published: (2025) -
Revisiting Cosine Similarity via Normalized ICA-transformed Embeddings
by: Yamagiwa, Hiroaki, et al.
Published: (2024)