Using Shapley interactions to understand how models use structure
Fuente:
arXiv
Saved in:
| Main Authors: | Singhvi, Divyansh, Misra, Diganta, Erkelens, Andrej, Jain, Raghav, Papadimitriou, Isabel, Saphra, Naomi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpreting the linear structure of vision-language model embedding spaces
by: Papadimitriou, Isabel, et al.
Published: (2025)
by: Papadimitriou, Isabel, et al.
Published: (2025)
Attribute Diversity Determines the Systematicity Gap in VQA
by: Berlot-Attwell, Ian, et al.
Published: (2023)
by: Berlot-Attwell, Ian, et al.
Published: (2023)
On the low-shot transferability of [V]-Mamba
by: Misra, Diganta, et al.
Published: (2024)
by: Misra, Diganta, et al.
Published: (2024)
ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling
by: Maina, Hernán, et al.
Published: (2025)
by: Maina, Hernán, et al.
Published: (2025)
ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter
by: Yuan, Zhengqing, et al.
Published: (2023)
by: Yuan, Zhengqing, et al.
Published: (2023)
ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding
by: Yang, Jianjiang, et al.
Published: (2025)
by: Yang, Jianjiang, et al.
Published: (2025)
FedEx-LoRA: Exact Aggregation for Federated and Efficient Fine-Tuning of Foundation Models
by: Singhal, Raghav, et al.
Published: (2024)
by: Singhal, Raghav, et al.
Published: (2024)
$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
by: Das, Trishanu, et al.
Published: (2025)
by: Das, Trishanu, et al.
Published: (2025)
mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs
by: Geigle, Gregor, et al.
Published: (2023)
by: Geigle, Gregor, et al.
Published: (2023)
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
(Almost) Free Modality Stitching of Foundation Models
by: Singh, Jaisidh, et al.
Published: (2025)
by: Singh, Jaisidh, et al.
Published: (2025)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
by: Faure, Gueter Josmy, et al.
Published: (2024)
by: Faure, Gueter Josmy, et al.
Published: (2024)
Uncovering the Hidden Cost of Model Compression
by: Misra, Diganta, et al.
Published: (2023)
by: Misra, Diganta, et al.
Published: (2023)
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables
by: Mathur, Suyash Vardhan, et al.
Published: (2024)
by: Mathur, Suyash Vardhan, et al.
Published: (2024)
Movie2Story: A framework for understanding videos and telling stories in the form of novel text
by: Li, Kangning, et al.
Published: (2024)
by: Li, Kangning, et al.
Published: (2024)
Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family
by: Chew, Oscar, et al.
Published: (2026)
by: Chew, Oscar, et al.
Published: (2026)
Link prediction Graph Neural Networks for structure recognition of Handwritten Mathematical Expressions
by: Nguyen, Cuong Tuan, et al.
Published: (2025)
by: Nguyen, Cuong Tuan, et al.
Published: (2025)
A General Framework for Inference-time Scaling and Steering of Diffusion Models
by: Singhal, Raghav, et al.
Published: (2025)
by: Singhal, Raghav, et al.
Published: (2025)
IMSAHLO: Integrating Multi-Scale Attention and Hybrid Loss Optimization Framework for Robust Neuronal Brain Cell Segmentation
by: Jain, Ujjwal, et al.
Published: (2026)
by: Jain, Ujjwal, et al.
Published: (2026)
Towards Deployable OCR models for Indic languages
by: Mathew, Minesh, et al.
Published: (2022)
by: Mathew, Minesh, et al.
Published: (2022)
TrafficLens: Multi-Camera Traffic Video Analysis Using LLMs
by: Arefeen, Md Adnan, et al.
Published: (2025)
by: Arefeen, Md Adnan, et al.
Published: (2025)
Cross-attention for State-based model RWKV-7
by: Xiao, Liu, et al.
Published: (2025)
by: Xiao, Liu, et al.
Published: (2025)
Image captioning for Brazilian Portuguese using GRIT model
by: de Alencar, Rafael Silva, et al.
Published: (2024)
by: de Alencar, Rafael Silva, et al.
Published: (2024)
Using Sign Language Production as Data Augmentation to enhance Sign Language Translation
by: Walsh, Harry, et al.
Published: (2025)
by: Walsh, Harry, et al.
Published: (2025)
Visual Confused Deputy: Exploiting and Defending Perception Failures in Computer-Using Agents
by: Liu, Xunzhuo, et al.
Published: (2026)
by: Liu, Xunzhuo, et al.
Published: (2026)
An Efficient Sign Language Translation Using Spatial Configuration and Motion Dynamics with LLMs
by: Hwang, Eui Jun, et al.
Published: (2024)
by: Hwang, Eui Jun, et al.
Published: (2024)
Using Multimodal Deep Neural Networks to Disentangle Language from Visual Aesthetics
by: Conwell, Colin, et al.
Published: (2024)
by: Conwell, Colin, et al.
Published: (2024)
Understanding Museum Exhibits using Vision-Language Reasoning
by: Balauca, Ada-Astrid, et al.
Published: (2024)
by: Balauca, Ada-Astrid, et al.
Published: (2024)
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
by: Beňová, Ivana, et al.
Published: (2024)
by: Beňová, Ivana, et al.
Published: (2024)
Developing an efficient corpus using Ensemble Data cleaning approach
by: Ahad, Md Taimur
Published: (2024)
by: Ahad, Md Taimur
Published: (2024)
Heart Disease Prediction using Case Based Reasoning (CBR)
by: Bhuiyan, Mohaiminul Islam, et al.
Published: (2025)
by: Bhuiyan, Mohaiminul Islam, et al.
Published: (2025)
Using Prompts to Guide Large Language Models in Imitating a Real Person's Language Style
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
Detecting Offensive Memes with Social Biases in Singapore Context Using Multimodal Large Language Models
by: Yuxuan, Cao, et al.
Published: (2025)
by: Yuxuan, Cao, et al.
Published: (2025)
Charting the Future: Using Chart Question-Answering for Scalable Evaluation of LLM-Driven Data Visualizations
by: Ford, James, et al.
Published: (2024)
by: Ford, James, et al.
Published: (2024)
NOLA: Compressing LoRA using Linear Combination of Random Basis
by: Koohpayegani, Soroush Abbasi, et al.
Published: (2023)
by: Koohpayegani, Soroush Abbasi, et al.
Published: (2023)
Open World Scene Graph Generation using Vision Language Models
by: Dutta, Amartya, et al.
Published: (2025)
by: Dutta, Amartya, et al.
Published: (2025)
Tailored Design of Audio-Visual Speech Recognition Models using Branchformers
by: Gimeno-Gómez, David, et al.
Published: (2024)
by: Gimeno-Gómez, David, et al.
Published: (2024)
DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images
by: Dalal, Dwip, et al.
Published: (2025)
by: Dalal, Dwip, et al.
Published: (2025)
Frame Sampling Strategies Matter: A Benchmark for small vision language models
by: Brkic, Marija, et al.
Published: (2025)
by: Brkic, Marija, et al.
Published: (2025)
Similar Items
-
Interpreting the linear structure of vision-language model embedding spaces
by: Papadimitriou, Isabel, et al.
Published: (2025) -
Attribute Diversity Determines the Systematicity Gap in VQA
by: Berlot-Attwell, Ian, et al.
Published: (2023) -
On the low-shot transferability of [V]-Mamba
by: Misra, Diganta, et al.
Published: (2024) -
ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling
by: Maina, Hernán, et al.
Published: (2025) -
ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter
by: Yuan, Zhengqing, et al.
Published: (2023)