Towards Interpreting Visual Information Processing in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Neo, Clement, Ong, Luke, Torr, Philip, Geva, Mor, Krueger, David, Barez, Fazl |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpreting Learned Feedback Patterns in Large Language Models
von: Marks, Luke, et al.
Veröffentlicht: (2023)
von: Marks, Luke, et al.
Veröffentlicht: (2023)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)
von: Lan, Michael, et al.
Veröffentlicht: (2023)
Understanding Addition and Subtraction in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
von: Marks, Luke, et al.
Veröffentlicht: (2024)
von: Marks, Luke, et al.
Veröffentlicht: (2024)
How Visual Representations Map to Language Feature Space in Multimodal LLMs
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
von: Lan, Michael, et al.
Veröffentlicht: (2024)
von: Lan, Michael, et al.
Veröffentlicht: (2024)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
von: Parmar, Mihir, et al.
Veröffentlicht: (2022)
von: Parmar, Mihir, et al.
Veröffentlicht: (2022)
VFusion3D: Learning Scalable 3D Generative Models from Video Diffusion Models
von: Han, Junlin, et al.
Veröffentlicht: (2024)
von: Han, Junlin, et al.
Veröffentlicht: (2024)
Interpreting Neurons in Deep Vision Networks with Language Models
von: Bai, Nicholas, et al.
Veröffentlicht: (2024)
von: Bai, Nicholas, et al.
Veröffentlicht: (2024)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
von: Oldfield, James, et al.
Veröffentlicht: (2025)
von: Oldfield, James, et al.
Veröffentlicht: (2025)
Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex
von: Grosbard, Idan Daniel, et al.
Veröffentlicht: (2026)
von: Grosbard, Idan Daniel, et al.
Veröffentlicht: (2026)
Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information
von: Chu, Xu, et al.
Veröffentlicht: (2025)
von: Chu, Xu, et al.
Veröffentlicht: (2025)
Attribute-based Visual Reprogramming for Vision-Language Models
von: Cai, Chengyi, et al.
Veröffentlicht: (2025)
von: Cai, Chengyi, et al.
Veröffentlicht: (2025)
Do Sparse Autoencoders Generalize? A Case Study of Answerability
von: Heindrich, Lovis, et al.
Veröffentlicht: (2025)
von: Heindrich, Lovis, et al.
Veröffentlicht: (2025)
Class-Discriminative Attention Maps for Vision Transformers
von: Brocki, Lennart, et al.
Veröffentlicht: (2023)
von: Brocki, Lennart, et al.
Veröffentlicht: (2023)
EchoPrime: A Multi-Video View-Informed Vision-Language Model for Comprehensive Echocardiography Interpretation
von: Vukadinovic, Milos, et al.
Veröffentlicht: (2024)
von: Vukadinovic, Milos, et al.
Veröffentlicht: (2024)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
von: Han, Junlin, et al.
Veröffentlicht: (2025)
von: Han, Junlin, et al.
Veröffentlicht: (2025)
Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
von: Jiang, Nick, et al.
Veröffentlicht: (2024)
von: Jiang, Nick, et al.
Veröffentlicht: (2024)
Improving Adversarial Transferability via Model Alignment
von: Ma, Avery, et al.
Veröffentlicht: (2023)
von: Ma, Avery, et al.
Veröffentlicht: (2023)
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
von: Rajabi, Navid, et al.
Veröffentlicht: (2023)
von: Rajabi, Navid, et al.
Veröffentlicht: (2023)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
VORD: Visual Ordinal Calibration for Mitigating Object Hallucinations in Large Vision-Language Models
von: Neo, Dexter, et al.
Veröffentlicht: (2024)
von: Neo, Dexter, et al.
Veröffentlicht: (2024)
Towards Understanding Multimodal Fine-Tuning: Spatial Features
von: Naghashyar, Lachin, et al.
Veröffentlicht: (2026)
von: Naghashyar, Lachin, et al.
Veröffentlicht: (2026)
Towards Compatible Fine-tuning for Vision-Language Model Updates
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
von: Berasi, Davide, et al.
Veröffentlicht: (2025)
von: Berasi, Davide, et al.
Veröffentlicht: (2025)
EasyARC: Evaluating Vision Language Models on True Visual Reasoning
von: Unsal, Mert, et al.
Veröffentlicht: (2025)
von: Unsal, Mert, et al.
Veröffentlicht: (2025)
Towards Improved Cervical Cancer Screening: Vision Transformer-Based Classification and Interpretability
von: Nguyen, Khoa Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Khoa Tuan, et al.
Veröffentlicht: (2025)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
von: Prabhu, Ameya, et al.
Veröffentlicht: (2024)
von: Prabhu, Ameya, et al.
Veröffentlicht: (2024)
IPO: Interpretable Prompt Optimization for Vision-Language Models
von: Du, Yingjun, et al.
Veröffentlicht: (2024)
von: Du, Yingjun, et al.
Veröffentlicht: (2024)
Unveiling the Visual Counting Bottleneck in Vision-Language Models
von: Pang, Xingzhou, et al.
Veröffentlicht: (2026)
von: Pang, Xingzhou, et al.
Veröffentlicht: (2026)
Towards Understanding How Knowledge Evolves in Large Vision-Language Models
von: Wang, Sudong, et al.
Veröffentlicht: (2025)
von: Wang, Sudong, et al.
Veröffentlicht: (2025)
Forecasting and Visualizing Air Quality from Sky Images with Vision-Language Models
von: Vahdatpour, Mohammad Saleh, et al.
Veröffentlicht: (2025)
von: Vahdatpour, Mohammad Saleh, et al.
Veröffentlicht: (2025)
Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models
von: Hou, Kuinan, et al.
Veröffentlicht: (2025)
von: Hou, Kuinan, et al.
Veröffentlicht: (2025)
Visual Text Meets Low-level Vision: A Comprehensive Survey on Visual Text Processing
von: Shu, Yan, et al.
Veröffentlicht: (2024)
von: Shu, Yan, et al.
Veröffentlicht: (2024)
Information Router for Mitigating Modality Dominance in Vision-Language Models
von: Kim, Seulgi, et al.
Veröffentlicht: (2026)
von: Kim, Seulgi, et al.
Veröffentlicht: (2026)
PVI: Plug-in Visual Injection for Vision-Language-Action Models
von: Zhang, Zezhou, et al.
Veröffentlicht: (2026)
von: Zhang, Zezhou, et al.
Veröffentlicht: (2026)
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models
von: Li, Xu, et al.
Veröffentlicht: (2024)
von: Li, Xu, et al.
Veröffentlicht: (2024)
Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
von: Hartsock, Iryna, et al.
Veröffentlicht: (2024)
von: Hartsock, Iryna, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Interpreting Learned Feedback Patterns in Large Language Models
von: Marks, Luke, et al.
Veröffentlicht: (2023) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023) -
Understanding Addition and Subtraction in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2024) -
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024) -
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
von: Marks, Luke, et al.
Veröffentlicht: (2024)