Mechanistically Interpreting Compression in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Elluru, Veeraraju, Singh, Arth, Aguero, Roberto, Agarwal, Ajay, Das, Debojyoti, Paul, Hreetam |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bias-Aware Machine Unlearning: Towards Fairer Vision Models via Controllable Forgetting
by: Aylapuram, Sai Siddhartha Chary, et al.
Published: (2025)
by: Aylapuram, Sai Siddhartha Chary, et al.
Published: (2025)
Re-Mask and Redirect: Exploiting Denoising Irreversibility in Diffusion Language Models
by: Singh, Arth
Published: (2026)
by: Singh, Arth
Published: (2026)
EMA Is Not All You Need: Mapping the Boundary Between Structure and Content in Recurrent Context
by: Singh, Arth
Published: (2026)
by: Singh, Arth
Published: (2026)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
by: Mahale, Ajay Pravin
Published: (2026)
by: Mahale, Ajay Pravin
Published: (2026)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
by: Che, Liwei, et al.
Published: (2026)
by: Che, Liwei, et al.
Published: (2026)
Mechanistic Interpretability of Emotion Inference in Large Language Models
by: Tak, Ala N., et al.
Published: (2025)
by: Tak, Ala N., et al.
Published: (2025)
Mechanistic Behavior Editing of Language Models
by: Singh, Joykirat, et al.
Published: (2024)
by: Singh, Joykirat, et al.
Published: (2024)
Mechanistic Interpretability of Socio-Political Frames in Language Models
by: Asghari, Hadi, et al.
Published: (2025)
by: Asghari, Hadi, et al.
Published: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
Challenges in Mechanistically Interpreting Model Representations
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
by: Żukowska, Nina, et al.
Published: (2026)
by: Żukowska, Nina, et al.
Published: (2026)
Beyond Protein Language Models: An Agentic LLM Framework for Mechanistic Enzyme Design
by: Jacob, Bruno, et al.
Published: (2025)
by: Jacob, Bruno, et al.
Published: (2025)
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
by: Winninger, Thomas, et al.
Published: (2025)
by: Winninger, Thomas, et al.
Published: (2025)
Supernova Event Dataset: Interpreting Large Language Models' Personality through Critical Event Analysis
by: Agarwal, Pranav, et al.
Published: (2025)
by: Agarwal, Pranav, et al.
Published: (2025)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024)
by: Kim, Geonhee, et al.
Published: (2024)
MALCDF: A Distributed Multi-Agent LLM Framework for Real-Time Cyber
by: Bhardwaj, Arth, et al.
Published: (2025)
by: Bhardwaj, Arth, et al.
Published: (2025)
Mechanistic Interpretability of LoRA-Adapted Language Models for Nuclear Reactor Safety Applications
by: Lee, Yoon Pyo
Published: (2025)
by: Lee, Yoon Pyo
Published: (2025)
Mechanistic Interpretability Needs Philosophy
by: Williams, Iwan, et al.
Published: (2025)
by: Williams, Iwan, et al.
Published: (2025)
Mechanistic Interpretability for AI Safety -- A Review
by: Bereska, Leonard, et al.
Published: (2024)
by: Bereska, Leonard, et al.
Published: (2024)
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
by: Rai, Daking, et al.
Published: (2024)
by: Rai, Daking, et al.
Published: (2024)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
by: Wang, Xu, et al.
Published: (2026)
by: Wang, Xu, et al.
Published: (2026)
Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models
by: Simbeck, Katharina, et al.
Published: (2025)
by: Simbeck, Katharina, et al.
Published: (2025)
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
by: Geiger, Atticus, et al.
Published: (2023)
by: Geiger, Atticus, et al.
Published: (2023)
Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes
by: Maghsoudi, Maryam, et al.
Published: (2026)
by: Maghsoudi, Maryam, et al.
Published: (2026)
reward-lens: A Mechanistic Interpretability Library for Reward Models
by: Nadaf, Mohammed Suhail B
Published: (2026)
by: Nadaf, Mohammed Suhail B
Published: (2026)
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
by: Lin, Zihao, et al.
Published: (2025)
by: Lin, Zihao, et al.
Published: (2025)
Mechanistic Interpretability in the Presence of Architectural Obfuscation
by: Florencio, Marcos, et al.
Published: (2025)
by: Florencio, Marcos, et al.
Published: (2025)
Auditing Disability Representation in Vision-Language Models
by: Panda, Srikant, et al.
Published: (2026)
by: Panda, Srikant, et al.
Published: (2026)
Beyond BeautifulSoup: Benchmarking LLM-Powered Web Scraping for Everyday Users
by: Bhardwaj, Arth, et al.
Published: (2026)
by: Bhardwaj, Arth, et al.
Published: (2026)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
by: Beniwal, Himanshu, et al.
Published: (2026)
by: Beniwal, Himanshu, et al.
Published: (2026)
The Security Threat of Compressed Projectors in Large Vision-Language Models
by: Zhang, Yudong, et al.
Published: (2025)
by: Zhang, Yudong, et al.
Published: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
by: Mishra, Anurag
Published: (2025)
by: Mishra, Anurag
Published: (2025)
Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models
by: Pawar, Pranav, et al.
Published: (2025)
by: Pawar, Pranav, et al.
Published: (2025)
On the Mechanistic Interpretability of Neural Networks for Causality in Bio-statistics
by: Conan, Jean-Baptiste A.
Published: (2025)
by: Conan, Jean-Baptiste A.
Published: (2025)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
by: He, Jesse, et al.
Published: (2026)
by: He, Jesse, et al.
Published: (2026)
RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models
by: Chen, Yuxuan, et al.
Published: (2025)
by: Chen, Yuxuan, et al.
Published: (2025)
Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective
by: Lee, Jae Hee, et al.
Published: (2025)
by: Lee, Jae Hee, et al.
Published: (2025)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
by: Seth, Ashish, et al.
Published: (2024)
by: Seth, Ashish, et al.
Published: (2024)
Similar Items
-
Bias-Aware Machine Unlearning: Towards Fairer Vision Models via Controllable Forgetting
by: Aylapuram, Sai Siddhartha Chary, et al.
Published: (2025) -
Re-Mask and Redirect: Exploiting Denoising Irreversibility in Diffusion Language Models
by: Singh, Arth
Published: (2026) -
EMA Is Not All You Need: Mapping the Boundary Between Structure and Content in Recurrent Context
by: Singh, Arth
Published: (2026) -
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
by: Mahale, Ajay Pravin
Published: (2026) -
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)