Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits
Fuente:
arXiv
Guardado en:
| Autores principales: | Mann, Logan, Saravanan, Ajit, Dave, Ishan, Shiromani, Shikhar, Ismail, Saadullah, Xia, Yi, Huang, Emily |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Linear Predictability of Attention Heads in Large Language Models
por: Shaikh, Khalid, et al.
Publicado: (2026)
por: Shaikh, Khalid, et al.
Publicado: (2026)
Focus Where It Matters: Graph Selective State Focused Attention Networks
por: Vashistha, Shikhar, et al.
Publicado: (2024)
por: Vashistha, Shikhar, et al.
Publicado: (2024)
The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders
por: Shiromani, Shikhar, et al.
Publicado: (2026)
por: Shiromani, Shikhar, et al.
Publicado: (2026)
COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
por: Sahay, Kenji, et al.
Publicado: (2025)
por: Sahay, Kenji, et al.
Publicado: (2025)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2026)
por: Beniwal, Himanshu, et al.
Publicado: (2026)
GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space
por: Shatwell, David G., et al.
Publicado: (2025)
por: Shatwell, David G., et al.
Publicado: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
por: Thomas, Rohan Subramanian, et al.
Publicado: (2026)
por: Thomas, Rohan Subramanian, et al.
Publicado: (2026)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
por: Che, Liwei, et al.
Publicado: (2026)
por: Che, Liwei, et al.
Publicado: (2026)
Predictive Behavioral Detection in Frozen Language Model Hidden States: Evidence for Pre-Surface Behavioral En- coding
por: Napolitano, Logan Matthew
Publicado: (2026)
por: Napolitano, Logan Matthew
Publicado: (2026)
Bridging Hidden States in Vision-Language Models
por: Fein-Ashley, Benjamin, et al.
Publicado: (2025)
por: Fein-Ashley, Benjamin, et al.
Publicado: (2025)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
por: Mitra, Chancharik, et al.
Publicado: (2025)
por: Mitra, Chancharik, et al.
Publicado: (2025)
Mechanistic Interpretability of Diffusion Models: Circuit-Level Analysis and Causal Validation
por: Roy, Dip
Publicado: (2025)
por: Roy, Dip
Publicado: (2025)
A finite difference scheme for two-dimensional singularly perturbed convection-diffusion problem with discontinuous source term
por: Shiromani, Ram, et al.
Publicado: (2024)
por: Shiromani, Ram, et al.
Publicado: (2024)
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
por: Żukowska, Nina, et al.
Publicado: (2026)
por: Żukowska, Nina, et al.
Publicado: (2026)
Rethinking Causal Mask Attention for Vision-Language Inference
por: Pei, Xiaohuan, et al.
Publicado: (2025)
por: Pei, Xiaohuan, et al.
Publicado: (2025)
Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
por: Kong, Injin, et al.
Publicado: (2026)
por: Kong, Injin, et al.
Publicado: (2026)
Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models
por: Pham, Minh Vu, et al.
Publicado: (2026)
por: Pham, Minh Vu, et al.
Publicado: (2026)
Cross-Modal Attention Analysis and Optimization in Vision-Language Models: A Study on Visual Reliability
por: Zhou, Lijie
Publicado: (2026)
por: Zhou, Lijie
Publicado: (2026)
Investigating Mechanisms for In-Context Vision Language Binding
por: Saravanan, Darshana, et al.
Publicado: (2025)
por: Saravanan, Darshana, et al.
Publicado: (2025)
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
por: Li, Qiming, et al.
Publicado: (2025)
por: Li, Qiming, et al.
Publicado: (2025)
Integrating Large Language Models with Network Optimization for Interactive and Explainable Supply Chain Planning: A Real-World Case Study
por: Venkatachalam, Saravanan
Publicado: (2025)
por: Venkatachalam, Saravanan
Publicado: (2025)
Certified Circuits: Stability Guarantees for Mechanistic Circuits
por: Anani, Alaa, et al.
Publicado: (2026)
por: Anani, Alaa, et al.
Publicado: (2026)
Where We Live Matters: Varied Disease Burden and Treatment Patterns in Psoriatic Arthritis in the United States
por: Brittany Banbury, et al.
Publicado: (2026)
por: Brittany Banbury, et al.
Publicado: (2026)
Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions
por: Ma, Nicole, et al.
Publicado: (2026)
por: Ma, Nicole, et al.
Publicado: (2026)
Mechanistic Circuit-Based Knowledge Editing in Large Language Models
por: Zhao, Tianyi, et al.
Publicado: (2026)
por: Zhao, Tianyi, et al.
Publicado: (2026)
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
por: Fioresi, Joseph, et al.
Publicado: (2025)
por: Fioresi, Joseph, et al.
Publicado: (2025)
ALBAR: Adversarial Learning approach to mitigate Biases in Action Recognition
por: Fioresi, Joseph, et al.
Publicado: (2025)
por: Fioresi, Joseph, et al.
Publicado: (2025)
Unmasking Hallucinations: A Causal Graph-Attention Perspective on Factual Reliability in Large Language Models
por: kurra, Sailesh kiran, et al.
Publicado: (2026)
por: kurra, Sailesh kiran, et al.
Publicado: (2026)
Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
por: Salla, Rohit Kumar, et al.
Publicado: (2025)
por: Salla, Rohit Kumar, et al.
Publicado: (2025)
HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
por: Jiang, Yilei, et al.
Publicado: (2025)
por: Jiang, Yilei, et al.
Publicado: (2025)
When and Where to Attack? Stage-wise Attention-Guided Adversarial Attack on Large Vision Language Models
por: Kwak, Jaehyun, et al.
Publicado: (2026)
por: Kwak, Jaehyun, et al.
Publicado: (2026)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
por: Song, Shezheng, et al.
Publicado: (2026)
por: Song, Shezheng, et al.
Publicado: (2026)
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
por: Ran-Milo, Yuval, et al.
Publicado: (2026)
por: Ran-Milo, Yuval, et al.
Publicado: (2026)
Hidden-Heavy Pentaquarks and Where to Find Them
por: Alasiri, Fareed, et al.
Publicado: (2025)
por: Alasiri, Fareed, et al.
Publicado: (2025)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
por: Kim, Geonhee, et al.
Publicado: (2024)
por: Kim, Geonhee, et al.
Publicado: (2024)
Mechanistically Interpreting Compression in Vision-Language Models
por: Elluru, Veeraraju, et al.
Publicado: (2026)
por: Elluru, Veeraraju, et al.
Publicado: (2026)
Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
por: Haskins, Reilly, et al.
Publicado: (2025)
por: Haskins, Reilly, et al.
Publicado: (2025)
On Mechanistic Circuits for Extractive Question-Answering
por: Basu, Samyadeep, et al.
Publicado: (2025)
por: Basu, Samyadeep, et al.
Publicado: (2025)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
por: Shukla, Shikhar
Publicado: (2026)
por: Shukla, Shikhar
Publicado: (2026)
Circular horizons: Pioneering sustainable paths in the oil and gas Industry's journey to net‐zero resilience
por: Shikhar Dua
Publicado: (2024)
por: Shikhar Dua
Publicado: (2024)
Ejemplares similares
-
Linear Predictability of Attention Heads in Large Language Models
por: Shaikh, Khalid, et al.
Publicado: (2026) -
Focus Where It Matters: Graph Selective State Focused Attention Networks
por: Vashistha, Shikhar, et al.
Publicado: (2024) -
The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders
por: Shiromani, Shikhar, et al.
Publicado: (2026) -
COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
por: Sahay, Kenji, et al.
Publicado: (2025) -
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2026)