A Concept-Based Explainability Framework for Large Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909409729314816 |
|---|---|
| author | Parekh, Jayneel Khayatan, Pegah Shukor, Mustafa Newson, Alasdair Cord, Matthieu |
| author_facet | Parekh, Jayneel Khayatan, Pegah Shukor, Mustafa Newson, Alasdair Cord, Matthieu |
| contents | Large multimodal models (LMMs) combine unimodal encoders and large language models (LLMs) to perform multimodal tasks. Despite recent advancements towards the interpretability of these models, understanding internal representations of LMMs remains largely a mystery. In this paper, we present a novel framework for the interpretation of LMMs. We propose a dictionary learning based approach, applied to the representation of tokens. The elements of the learned dictionary correspond to our proposed concepts. We show that these concepts are well semantically grounded in both vision and text. Thus we refer to these as ``multi-modal concepts''. We qualitatively and quantitatively evaluate the results of the learnt concepts. We show that the extracted multimodal concepts are useful to interpret representations of test samples. Finally, we evaluate the disentanglement between different concepts and the quality of grounding concepts visually and textually. Our code is publicly available at https://github.com/mshukor/xl-vlms |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_08074 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Concept-Based Explainability Framework for Large Multimodal Models Parekh, Jayneel Khayatan, Pegah Shukor, Mustafa Newson, Alasdair Cord, Matthieu Machine Learning Artificial Intelligence Computation and Language Computer Vision and Pattern Recognition Large multimodal models (LMMs) combine unimodal encoders and large language models (LLMs) to perform multimodal tasks. Despite recent advancements towards the interpretability of these models, understanding internal representations of LMMs remains largely a mystery. In this paper, we present a novel framework for the interpretation of LMMs. We propose a dictionary learning based approach, applied to the representation of tokens. The elements of the learned dictionary correspond to our proposed concepts. We show that these concepts are well semantically grounded in both vision and text. Thus we refer to these as ``multi-modal concepts''. We qualitatively and quantitatively evaluate the results of the learnt concepts. We show that the extracted multimodal concepts are useful to interpret representations of test samples. Finally, we evaluate the disentanglement between different concepts and the quality of grounding concepts visually and textually. Our code is publicly available at https://github.com/mshukor/xl-vlms |
| title | A Concept-Based Explainability Framework for Large Multimodal Models |
| topic | Machine Learning Artificial Intelligence Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2406.08074 |