Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Yanbei, Ma, Xueqi, Liu, Shu, Erfani, Sarah Monazam, Liu, Tongliang, Bailey, James, Lau, Jey Han, Ehinger, Krista A. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
by: Ma, Xueqi, et al.
Published: (2025)
by: Ma, Xueqi, et al.
Published: (2025)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026)
by: Ma, Xueqi, et al.
Published: (2026)
Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
by: Ma, Xueqi, et al.
Published: (2025)
by: Ma, Xueqi, et al.
Published: (2025)
KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph
by: Jiang, Yanbei, et al.
Published: (2024)
by: Jiang, Yanbei, et al.
Published: (2024)
PROPA: Toward Process-level Optimization in Visual Reasoning via Reinforcement Learning
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
Coarse-to-Fine Open-Set Graph Node Classification with Large Language Models
by: Ma, Xueqi, et al.
Published: (2025)
by: Ma, Xueqi, et al.
Published: (2025)
Unlearnable Examples For Time Series
by: Jiang, Yujing, et al.
Published: (2024)
by: Jiang, Yujing, et al.
Published: (2024)
End-to-End Anti-Backdoor Learning on Images and Time Series
by: Jiang, Yujing, et al.
Published: (2024)
by: Jiang, Yujing, et al.
Published: (2024)
Open-World Amodal Appearance Completion
by: Ao, Jiayang, et al.
Published: (2024)
by: Ao, Jiayang, et al.
Published: (2024)
Efficient Neural Implicit Representation for 3D Human Reconstruction
by: Huang, Zexu, et al.
Published: (2024)
by: Huang, Zexu, et al.
Published: (2024)
Perceiving Longer Sequences With Bi-Directional Cross-Attention Transformers
by: Hiller, Markus, et al.
Published: (2024)
by: Hiller, Markus, et al.
Published: (2024)
Generalized Planning for the Abstraction and Reasoning Corpus
by: Lei, Chao, et al.
Published: (2024)
by: Lei, Chao, et al.
Published: (2024)
LDReg: Local Dimensionality Regularized Self-Supervised Learning
by: Huang, Hanxun, et al.
Published: (2024)
by: Huang, Hanxun, et al.
Published: (2024)
Evaluating Evidence Attribution in Generated Fact Checking Explanations
by: Xing, Rui, et al.
Published: (2024)
by: Xing, Rui, et al.
Published: (2024)
Understanding the Geospatial Reasoning Capabilities of LLMs: A Trajectory Recovery Perspective
by: Truong, Thinh Hung, et al.
Published: (2025)
by: Truong, Thinh Hung, et al.
Published: (2025)
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
by: Jiang, Yanbei, et al.
Published: (2026)
by: Jiang, Yanbei, et al.
Published: (2026)
Decomposed Opinion Summarization with Verified Aspect-Aware Modules
by: Li, Miao, et al.
Published: (2025)
by: Li, Miao, et al.
Published: (2025)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
by: Lei, Chao, et al.
Published: (2025)
by: Lei, Chao, et al.
Published: (2025)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
by: Li, Jinhao, et al.
Published: (2024)
by: Li, Jinhao, et al.
Published: (2024)
Sequential Amodal Segmentation via Cumulative Occlusion Learning
by: Ao, Jiayang, et al.
Published: (2024)
by: Ao, Jiayang, et al.
Published: (2024)
State-Based Disassembly Planning
by: Lei, Chao, et al.
Published: (2025)
by: Lei, Chao, et al.
Published: (2025)
Improving Denoising Diffusion Models via Simultaneous Estimation of Image and Noise
by: Zhang, Zhenkai, et al.
Published: (2023)
by: Zhang, Zhenkai, et al.
Published: (2023)
Factual Dialogue Summarization via Learning from Large Language Models
by: Zhu, Rongxin, et al.
Published: (2024)
by: Zhu, Rongxin, et al.
Published: (2024)
CMA-R:Causal Mediation Analysis for Explaining Rumour Detection
by: Tian, Lin, et al.
Published: (2024)
by: Tian, Lin, et al.
Published: (2024)
MoDEM: Mixture of Domain Expert Models
by: Simonds, Toby, et al.
Published: (2024)
by: Simonds, Toby, et al.
Published: (2024)
Interaction Matters: An Evaluation Framework for Interactive Dialogue Assessment on English Second Language Conversations
by: Gao, Rena, et al.
Published: (2024)
by: Gao, Rena, et al.
Published: (2024)
REL: Working out is all you need
by: Simonds, Toby, et al.
Published: (2024)
by: Simonds, Toby, et al.
Published: (2024)
WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks
by: Shetty, Anudeex, et al.
Published: (2024)
by: Shetty, Anudeex, et al.
Published: (2024)
Beyond Seen Data: Improving KBQA Generalization Through Schema-Guided Logical Form Generation
by: Gao, Shengxiang, et al.
Published: (2025)
by: Gao, Shengxiang, et al.
Published: (2025)
A Sentiment Consolidation Framework for Meta-Review Generation
by: Li, Miao, et al.
Published: (2024)
by: Li, Miao, et al.
Published: (2024)
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
Detecting Backdoor Samples in Contrastive Language Image Pretraining
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
GPSBench: Do Large Language Models Understand GPS Coordinates?
by: Truong, Thinh Hung, et al.
Published: (2026)
by: Truong, Thinh Hung, et al.
Published: (2026)
WHoW: A Cross-domain Approach for Analysing Conversation Moderation
by: Chen, Ming-Bin, et al.
Published: (2024)
by: Chen, Ming-Bin, et al.
Published: (2024)
Context Volume Drives Performance: Tackling Domain Shift in Extremely Low-Resource Translation via RAG
by: Setiawan, David Samuel, et al.
Published: (2026)
by: Setiawan, David Samuel, et al.
Published: (2026)
CIG: Measuring Conversational Information Gain in Deliberative Dialogues with Semantic Memory Dynamics
by: Chen, Ming-Bin, et al.
Published: (2026)
by: Chen, Ming-Bin, et al.
Published: (2026)
Exploring Weak-to-Strong Generalization for CLIP-based Classification
by: Li, Jinhao, et al.
Published: (2025)
by: Li, Jinhao, et al.
Published: (2025)
Planning-Driven Programming: A Large Language Model Programming Workflow
by: Lei, Chao, et al.
Published: (2024)
by: Lei, Chao, et al.
Published: (2024)
On the Interplay between Human Label Variation and Model Fairness
by: Kurniawan, Kemal, et al.
Published: (2025)
by: Kurniawan, Kemal, et al.
Published: (2025)
Similar Items
-
Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
by: Ma, Xueqi, et al.
Published: (2025) -
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026) -
Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
by: Ma, Xueqi, et al.
Published: (2025) -
KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph
by: Jiang, Yanbei, et al.
Published: (2024) -
PROPA: Toward Process-level Optimization in Visual Reasoning via Reinforcement Learning
by: Jiang, Yanbei, et al.
Published: (2025)