Interpreting Neurons in Deep Vision Networks with Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Nicholas, Iyer, Rahul A., Oikarinen, Tuomas, Kulkarni, Akshay, Weng, Tsui-Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Linear Explanations for Individual Neurons
by: Oikarinen, Tuomas, et al.
Published: (2024)
by: Oikarinen, Tuomas, et al.
Published: (2024)
Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
by: Oikarinen, Tuomas, et al.
Published: (2025)
by: Oikarinen, Tuomas, et al.
Published: (2025)
Interpretable Generative Models through Post-hoc Concept Bottlenecks
by: Kulkarni, Akshay, et al.
Published: (2025)
by: Kulkarni, Akshay, et al.
Published: (2025)
CI-CBM: Class-Incremental Concept Bottleneck Model for Interpretable Continual Learning
by: Javadi, Amirhosein, et al.
Published: (2026)
by: Javadi, Amirhosein, et al.
Published: (2026)
Interpretability-Guided Test-Time Adversarial Defense
by: Kulkarni, Akshay, et al.
Published: (2024)
by: Kulkarni, Akshay, et al.
Published: (2024)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
by: Kulkarni, Akshay, et al.
Published: (2025)
by: Kulkarni, Akshay, et al.
Published: (2025)
VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance
by: Srivastava, Divyansh, et al.
Published: (2024)
by: Srivastava, Divyansh, et al.
Published: (2024)
Crafting Large Language Models for Enhanced Interpretability
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Evaluating Neuron Explanations: A Unified Framework with Sanity Checks
by: Oikarinen, Tuomas, et al.
Published: (2025)
by: Oikarinen, Tuomas, et al.
Published: (2025)
Medical Vision Language Models as Policies for Robotic Surgery
by: Muppidi, Akshay, et al.
Published: (2025)
by: Muppidi, Akshay, et al.
Published: (2025)
RAT: Boosting Misclassification Detection Ability without Extra Data
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Concept Bottleneck Large Language Models
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images
by: Flora, James, et al.
Published: (2026)
by: Flora, James, et al.
Published: (2026)
Provably Robust Conformal Prediction with Improved Efficiency
by: Yan, Ge, et al.
Published: (2024)
by: Yan, Ge, et al.
Published: (2024)
Hierarchical Invariance for Robust and Interpretable Vision Tasks at Larger Scales
by: Qi, Shuren, et al.
Published: (2024)
by: Qi, Shuren, et al.
Published: (2024)
Towards Interpreting Visual Information Processing in Vision-Language Models
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Cluster Paths: Navigating Interpretability in Neural Networks
by: Kroeger, Nicholas M., et al.
Published: (2025)
by: Kroeger, Nicholas M., et al.
Published: (2025)
Improved Object-Based Style Transfer with Single Deep Network
by: Kulkarni, Harshmohan, et al.
Published: (2024)
by: Kulkarni, Harshmohan, et al.
Published: (2024)
IPO: Interpretable Prompt Optimization for Vision-Language Models
by: Du, Yingjun, et al.
Published: (2024)
by: Du, Yingjun, et al.
Published: (2024)
On the Road with 16 Neurons: Mental Imagery with Bio-inspired Deep Neural Networks
by: Plebe, Alice, et al.
Published: (2020)
by: Plebe, Alice, et al.
Published: (2020)
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
by: Jiang, Nick, et al.
Published: (2024)
by: Jiang, Nick, et al.
Published: (2024)
MI CAM: Mutual Information Weighted Activation Mapping for Causal Visual Explanations of Convolutional Neural Networks
by: Iyer, Ram S, et al.
Published: (2025)
by: Iyer, Ram S, et al.
Published: (2025)
Discretized Quadratic Integrate-and-Fire Neuron Model for Deep Spiking Neural Networks
by: Jahns, Eric, et al.
Published: (2025)
by: Jahns, Eric, et al.
Published: (2025)
OSSCAR: One-Shot Structured Pruning in Vision and Language Models with Combinatorial Optimization
by: Meng, Xiang, et al.
Published: (2024)
by: Meng, Xiang, et al.
Published: (2024)
CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
by: Zhang, Mingyuan, et al.
Published: (2025)
by: Zhang, Mingyuan, et al.
Published: (2025)
Neutral-Reference Prompting for Vision-Language Models
by: Tian, Senmao, et al.
Published: (2026)
by: Tian, Senmao, et al.
Published: (2026)
Perturbation on Feature Coalition: Towards Interpretable Deep Neural Networks
by: Hu, Xuran, et al.
Published: (2024)
by: Hu, Xuran, et al.
Published: (2024)
A Reasoning-Enabled Vision-Language Foundation Model for Chest X-ray Interpretation
by: Zhang, Yabin, et al.
Published: (2026)
by: Zhang, Yabin, et al.
Published: (2026)
Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
EchoPrime: A Multi-Video View-Informed Vision-Language Model for Comprehensive Echocardiography Interpretation
by: Vukadinovic, Milos, et al.
Published: (2024)
by: Vukadinovic, Milos, et al.
Published: (2024)
FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing
by: Corley, Isaac, et al.
Published: (2025)
by: Corley, Isaac, et al.
Published: (2025)
Patronus: Interpretable Diffusion Models with Prototypes
by: Weng, Nina, et al.
Published: (2025)
by: Weng, Nina, et al.
Published: (2025)
Quantized and Interpretable Learning Scheme for Deep Neural Networks in Classification Task
by: Maleki, Alireza, et al.
Published: (2024)
by: Maleki, Alireza, et al.
Published: (2024)
Explaining Deep Convolutional Neural Networks for Image Classification by Evolving Local Interpretable Model-agnostic Explanations
by: Wang, Bin, et al.
Published: (2022)
by: Wang, Bin, et al.
Published: (2022)
Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
by: Xing, Shuo, et al.
Published: (2025)
by: Xing, Shuo, et al.
Published: (2025)
Learning without Forgetting for Vision-Language Models
by: Zhou, Da-Wei, et al.
Published: (2023)
by: Zhou, Da-Wei, et al.
Published: (2023)
VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set
by: Shen, Shufan, et al.
Published: (2025)
by: Shen, Shufan, et al.
Published: (2025)
Interpretable Diffusion Models with B-cos Networks
by: Bernold, Nicola, et al.
Published: (2025)
by: Bernold, Nicola, et al.
Published: (2025)
Similar Items
-
Linear Explanations for Individual Neurons
by: Oikarinen, Tuomas, et al.
Published: (2024) -
Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
by: Oikarinen, Tuomas, et al.
Published: (2025) -
Interpretable Generative Models through Post-hoc Concept Bottlenecks
by: Kulkarni, Akshay, et al.
Published: (2025) -
CI-CBM: Class-Incremental Concept Bottleneck Model for Interpretable Continual Learning
by: Javadi, Amirhosein, et al.
Published: (2026) -
Interpretability-Guided Test-Time Adversarial Defense
by: Kulkarni, Akshay, et al.
Published: (2024)