Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits
Fuente:
arXiv
Salvato in:
| Autori principali: | Mann, Logan, Saravanan, Ajit, Dave, Ishan, Shiromani, Shikhar, Ismail, Saadullah, Xia, Yi, Huang, Emily |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Linear Predictability of Attention Heads in Large Language Models
di: Shaikh, Khalid, et al.
Pubblicazione: (2026)
di: Shaikh, Khalid, et al.
Pubblicazione: (2026)
Focus Where It Matters: Graph Selective State Focused Attention Networks
di: Vashistha, Shikhar, et al.
Pubblicazione: (2024)
di: Vashistha, Shikhar, et al.
Pubblicazione: (2024)
The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders
di: Shiromani, Shikhar, et al.
Pubblicazione: (2026)
di: Shiromani, Shikhar, et al.
Pubblicazione: (2026)
COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
di: Sahay, Kenji, et al.
Pubblicazione: (2025)
di: Sahay, Kenji, et al.
Pubblicazione: (2025)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
di: Beniwal, Himanshu, et al.
Pubblicazione: (2026)
di: Beniwal, Himanshu, et al.
Pubblicazione: (2026)
GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space
di: Shatwell, David G., et al.
Pubblicazione: (2025)
di: Shatwell, David G., et al.
Pubblicazione: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026)
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
di: Che, Liwei, et al.
Pubblicazione: (2026)
di: Che, Liwei, et al.
Pubblicazione: (2026)
Predictive Behavioral Detection in Frozen Language Model Hidden States: Evidence for Pre-Surface Behavioral En- coding
di: Napolitano, Logan Matthew
Pubblicazione: (2026)
di: Napolitano, Logan Matthew
Pubblicazione: (2026)
Bridging Hidden States in Vision-Language Models
di: Fein-Ashley, Benjamin, et al.
Pubblicazione: (2025)
di: Fein-Ashley, Benjamin, et al.
Pubblicazione: (2025)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
di: Mitra, Chancharik, et al.
Pubblicazione: (2025)
di: Mitra, Chancharik, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of Diffusion Models: Circuit-Level Analysis and Causal Validation
di: Roy, Dip
Pubblicazione: (2025)
di: Roy, Dip
Pubblicazione: (2025)
A finite difference scheme for two-dimensional singularly perturbed convection-diffusion problem with discontinuous source term
di: Shiromani, Ram, et al.
Pubblicazione: (2024)
di: Shiromani, Ram, et al.
Pubblicazione: (2024)
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
di: Żukowska, Nina, et al.
Pubblicazione: (2026)
di: Żukowska, Nina, et al.
Pubblicazione: (2026)
Rethinking Causal Mask Attention for Vision-Language Inference
di: Pei, Xiaohuan, et al.
Pubblicazione: (2025)
di: Pei, Xiaohuan, et al.
Pubblicazione: (2025)
Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
di: Kong, Injin, et al.
Pubblicazione: (2026)
di: Kong, Injin, et al.
Pubblicazione: (2026)
Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models
di: Pham, Minh Vu, et al.
Pubblicazione: (2026)
di: Pham, Minh Vu, et al.
Pubblicazione: (2026)
Cross-Modal Attention Analysis and Optimization in Vision-Language Models: A Study on Visual Reliability
di: Zhou, Lijie
Pubblicazione: (2026)
di: Zhou, Lijie
Pubblicazione: (2026)
Investigating Mechanisms for In-Context Vision Language Binding
di: Saravanan, Darshana, et al.
Pubblicazione: (2025)
di: Saravanan, Darshana, et al.
Pubblicazione: (2025)
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
di: Li, Qiming, et al.
Pubblicazione: (2025)
di: Li, Qiming, et al.
Pubblicazione: (2025)
Integrating Large Language Models with Network Optimization for Interactive and Explainable Supply Chain Planning: A Real-World Case Study
di: Venkatachalam, Saravanan
Pubblicazione: (2025)
di: Venkatachalam, Saravanan
Pubblicazione: (2025)
Certified Circuits: Stability Guarantees for Mechanistic Circuits
di: Anani, Alaa, et al.
Pubblicazione: (2026)
di: Anani, Alaa, et al.
Pubblicazione: (2026)
Where We Live Matters: Varied Disease Burden and Treatment Patterns in Psoriatic Arthritis in the United States
di: Brittany Banbury, et al.
Pubblicazione: (2026)
di: Brittany Banbury, et al.
Pubblicazione: (2026)
Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions
di: Ma, Nicole, et al.
Pubblicazione: (2026)
di: Ma, Nicole, et al.
Pubblicazione: (2026)
Mechanistic Circuit-Based Knowledge Editing in Large Language Models
di: Zhao, Tianyi, et al.
Pubblicazione: (2026)
di: Zhao, Tianyi, et al.
Pubblicazione: (2026)
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
di: Fioresi, Joseph, et al.
Pubblicazione: (2025)
di: Fioresi, Joseph, et al.
Pubblicazione: (2025)
ALBAR: Adversarial Learning approach to mitigate Biases in Action Recognition
di: Fioresi, Joseph, et al.
Pubblicazione: (2025)
di: Fioresi, Joseph, et al.
Pubblicazione: (2025)
Unmasking Hallucinations: A Causal Graph-Attention Perspective on Factual Reliability in Large Language Models
di: kurra, Sailesh kiran, et al.
Pubblicazione: (2026)
di: kurra, Sailesh kiran, et al.
Pubblicazione: (2026)
HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
di: Jiang, Yilei, et al.
Pubblicazione: (2025)
di: Jiang, Yilei, et al.
Pubblicazione: (2025)
Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
di: Salla, Rohit Kumar, et al.
Pubblicazione: (2025)
di: Salla, Rohit Kumar, et al.
Pubblicazione: (2025)
When and Where to Attack? Stage-wise Attention-Guided Adversarial Attack on Large Vision Language Models
di: Kwak, Jaehyun, et al.
Pubblicazione: (2026)
di: Kwak, Jaehyun, et al.
Pubblicazione: (2026)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
di: Song, Shezheng, et al.
Pubblicazione: (2026)
di: Song, Shezheng, et al.
Pubblicazione: (2026)
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
di: Ran-Milo, Yuval, et al.
Pubblicazione: (2026)
di: Ran-Milo, Yuval, et al.
Pubblicazione: (2026)
Hidden-Heavy Pentaquarks and Where to Find Them
di: Alasiri, Fareed, et al.
Pubblicazione: (2025)
di: Alasiri, Fareed, et al.
Pubblicazione: (2025)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
Mechanistically Interpreting Compression in Vision-Language Models
di: Elluru, Veeraraju, et al.
Pubblicazione: (2026)
di: Elluru, Veeraraju, et al.
Pubblicazione: (2026)
Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
di: Haskins, Reilly, et al.
Pubblicazione: (2025)
di: Haskins, Reilly, et al.
Pubblicazione: (2025)
On Mechanistic Circuits for Extractive Question-Answering
di: Basu, Samyadeep, et al.
Pubblicazione: (2025)
di: Basu, Samyadeep, et al.
Pubblicazione: (2025)
Underspecification in Language Modeling Tasks: A Causality-Informed Study of Gendered Pronoun Resolution
di: McMilin, Emily
Pubblicazione: (2022)
di: McMilin, Emily
Pubblicazione: (2022)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
di: Shukla, Shikhar
Pubblicazione: (2026)
di: Shukla, Shikhar
Pubblicazione: (2026)
Documenti analoghi
-
Linear Predictability of Attention Heads in Large Language Models
di: Shaikh, Khalid, et al.
Pubblicazione: (2026) -
Focus Where It Matters: Graph Selective State Focused Attention Networks
di: Vashistha, Shikhar, et al.
Pubblicazione: (2024) -
The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders
di: Shiromani, Shikhar, et al.
Pubblicazione: (2026) -
COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
di: Sahay, Kenji, et al.
Pubblicazione: (2025) -
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
di: Beniwal, Himanshu, et al.
Pubblicazione: (2026)