What Do VLMs NOTICE? A Mechanistic Interpretability Pipeline for Gaussian-Noise-free Text-Image Corruption and Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Golovanevsky, Michal, Rudman, William, Palit, Vedant, Singh, Ritambhara, Eickhoff, Carsten |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind
por: Rudman, William, et al.
Publicado: (2025)
por: Rudman, William, et al.
Publicado: (2025)
PiCME: Pipeline for Contrastive Modality Evaluation and Encoding in the MIMIC Dataset
por: Golovanevsky, Michal, et al.
Publicado: (2025)
por: Golovanevsky, Michal, et al.
Publicado: (2025)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
por: Rudman, William, et al.
Publicado: (2026)
por: Rudman, William, et al.
Publicado: (2026)
Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts
por: Golovanevsky, Michal, et al.
Publicado: (2025)
por: Golovanevsky, Michal, et al.
Publicado: (2025)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
por: Nemitz, Jonathan, et al.
Publicado: (2026)
por: Nemitz, Jonathan, et al.
Publicado: (2026)
Is There Knowledge Left to Extract? Evidence of Fragility in Medically Fine-Tuned Vision-Language Models
por: McLaughlin, Oliver, et al.
Publicado: (2026)
por: McLaughlin, Oliver, et al.
Publicado: (2026)
Stable Anisotropic Regularization
por: Rudman, William, et al.
Publicado: (2023)
por: Rudman, William, et al.
Publicado: (2023)
Outlier Dimensions Encode Task-Specific Knowledge
por: Rudman, William, et al.
Publicado: (2023)
por: Rudman, William, et al.
Publicado: (2023)
TRIM: Achieving Extreme Sparsity with Targeted Row-wise Iterative Metric-driven Pruning
por: Beck, Florentin, et al.
Publicado: (2025)
por: Beck, Florentin, et al.
Publicado: (2025)
One-Versus-Others Attention: Scalable Multimodal Integration for Biomedical Data
por: Golovanevsky, Michal, et al.
Publicado: (2023)
por: Golovanevsky, Michal, et al.
Publicado: (2023)
Retrieval Augmented Zero-Shot Text Classification
por: Abdullahi, Tassallah, et al.
Publicado: (2024)
por: Abdullahi, Tassallah, et al.
Publicado: (2024)
Less is More: Label-Guided Summarization of Procedural and Instructional Videos
por: Rajpal, Shreya, et al.
Publicado: (2026)
por: Rajpal, Shreya, et al.
Publicado: (2026)
APP: Accelerated Path Patching with Task-Specific Pruning
por: Andersen, Frauke, et al.
Publicado: (2025)
por: Andersen, Frauke, et al.
Publicado: (2025)
From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures
por: Rottach, Florian, et al.
Publicado: (2025)
por: Rottach, Florian, et al.
Publicado: (2025)
UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages
por: Abdullahi, Tassallah, et al.
Publicado: (2026)
por: Abdullahi, Tassallah, et al.
Publicado: (2026)
Paths Not Taken: Understanding and Mending the Multilingual Factual Recall Pipeline
por: Lu, Meng, et al.
Publicado: (2025)
por: Lu, Meng, et al.
Publicado: (2025)
K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction
por: Abdullahi, Tassallah, et al.
Publicado: (2025)
por: Abdullahi, Tassallah, et al.
Publicado: (2025)
The Persona Paradox: Medical Personas as Behavioral Priors in Clinical Language Models
por: Abdullahi, Tassallah, et al.
Publicado: (2026)
por: Abdullahi, Tassallah, et al.
Publicado: (2026)
CroCoSum: A Benchmark Dataset for Cross-Lingual Code-Switched Summarization
por: Zhang, Ruochen, et al.
Publicado: (2023)
por: Zhang, Ruochen, et al.
Publicado: (2023)
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
por: Braun, Joschka, et al.
Publicado: (2025)
por: Braun, Joschka, et al.
Publicado: (2025)
Adaptive Federated Learning Defences via Trust-Aware Deep Q-Networks
por: Palit, Vedant
Publicado: (2025)
por: Palit, Vedant
Publicado: (2025)
Talking Heads: Understanding Inter-layer Communication in Transformer Language Models
por: Merullo, Jack, et al.
Publicado: (2024)
por: Merullo, Jack, et al.
Publicado: (2024)
Circuit Component Reuse Across Tasks in Transformer Language Models
por: Merullo, Jack, et al.
Publicado: (2023)
por: Merullo, Jack, et al.
Publicado: (2023)
Language Models Implement Simple Word2Vec-style Vector Arithmetic
por: Merullo, Jack, et al.
Publicado: (2023)
por: Merullo, Jack, et al.
Publicado: (2023)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
por: Toker, Michael, et al.
Publicado: (2024)
por: Toker, Michael, et al.
Publicado: (2024)
MechIR: A Mechanistic Interpretability Framework for Information Retrieval
por: Parry, Andrew, et al.
Publicado: (2025)
por: Parry, Andrew, et al.
Publicado: (2025)
PubMed-Ophtha: An open resource for training ophthalmology vision-language models on scientific literature
por: Hallitschke, Verena Jasmin, et al.
Publicado: (2026)
por: Hallitschke, Verena Jasmin, et al.
Publicado: (2026)
A Survey on LLM-Assisted Clinical Trial Recruitment
por: Ghosh, Shrestha, et al.
Publicado: (2025)
por: Ghosh, Shrestha, et al.
Publicado: (2025)
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
por: Mohammadi, Seyedali, et al.
Publicado: (2024)
por: Mohammadi, Seyedali, et al.
Publicado: (2024)
Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
por: Singh, Anshul, et al.
Publicado: (2025)
por: Singh, Anshul, et al.
Publicado: (2025)
Do Images Speak Louder than Words? Investigating the Effect of Textual Misinformation in VLMs
por: Zhang, Chi, et al.
Publicado: (2026)
por: Zhang, Chi, et al.
Publicado: (2026)
Enhancing Retrieval-Augmented Generation: A Study of Best Practices
por: Li, Siran, et al.
Publicado: (2025)
por: Li, Siran, et al.
Publicado: (2025)
Large Language Models for Mental Health Diagnostic Assessments: Exploring The Potential of Large Language Models for Assisting with Mental Health Diagnostic Assessments -- The Depression and Anxiety Case
por: Roy, Kaushik, et al.
Publicado: (2025)
por: Roy, Kaushik, et al.
Publicado: (2025)
The Same But Different: Structural Similarities and Differences in Multilingual Language Modeling
por: Zhang, Ruochen, et al.
Publicado: (2024)
por: Zhang, Ruochen, et al.
Publicado: (2024)
Mechanistic Interpretability Needs Philosophy
por: Williams, Iwan, et al.
Publicado: (2025)
por: Williams, Iwan, et al.
Publicado: (2025)
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
por: Im, Shawn, et al.
Publicado: (2026)
por: Im, Shawn, et al.
Publicado: (2026)
Are VLMs Really Blind
por: Singh, Ayush, et al.
Publicado: (2024)
por: Singh, Ayush, et al.
Publicado: (2024)
Towards Reliable and Interpretable Document Question Answering via VLMs
por: Chen, Alessio, et al.
Publicado: (2025)
por: Chen, Alessio, et al.
Publicado: (2025)
Evaluating Search System Explainability with Psychometrics and Crowdsourcing
por: Chen, Catherine, et al.
Publicado: (2022)
por: Chen, Catherine, et al.
Publicado: (2022)
When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
por: Zhou, Xinyu, et al.
Publicado: (2026)
por: Zhou, Xinyu, et al.
Publicado: (2026)
Ejemplares similares
-
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind
por: Rudman, William, et al.
Publicado: (2025) -
PiCME: Pipeline for Contrastive Modality Evaluation and Encoding in the MIMIC Dataset
por: Golovanevsky, Michal, et al.
Publicado: (2025) -
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
por: Rudman, William, et al.
Publicado: (2026) -
Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts
por: Golovanevsky, Michal, et al.
Publicado: (2025) -
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
por: Nemitz, Jonathan, et al.
Publicado: (2026)