A Concept-Based Explainability Framework for Large Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Parekh, Jayneel, Khayatan, Pegah, Shukor, Mustafa, Newson, Alasdair, Cord, Matthieu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909409729314816
author Parekh, Jayneel
Khayatan, Pegah
Shukor, Mustafa
Newson, Alasdair
Cord, Matthieu
author_facet Parekh, Jayneel
Khayatan, Pegah
Shukor, Mustafa
Newson, Alasdair
Cord, Matthieu
contents Large multimodal models (LMMs) combine unimodal encoders and large language models (LLMs) to perform multimodal tasks. Despite recent advancements towards the interpretability of these models, understanding internal representations of LMMs remains largely a mystery. In this paper, we present a novel framework for the interpretation of LMMs. We propose a dictionary learning based approach, applied to the representation of tokens. The elements of the learned dictionary correspond to our proposed concepts. We show that these concepts are well semantically grounded in both vision and text. Thus we refer to these as ``multi-modal concepts''. We qualitatively and quantitatively evaluate the results of the learnt concepts. We show that the extracted multimodal concepts are useful to interpret representations of test samples. Finally, we evaluate the disentanglement between different concepts and the quality of grounding concepts visually and textually. Our code is publicly available at https://github.com/mshukor/xl-vlms
format Preprint
id arxiv_https___arxiv_org_abs_2406_08074
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Concept-Based Explainability Framework for Large Multimodal Models
Parekh, Jayneel
Khayatan, Pegah
Shukor, Mustafa
Newson, Alasdair
Cord, Matthieu
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Large multimodal models (LMMs) combine unimodal encoders and large language models (LLMs) to perform multimodal tasks. Despite recent advancements towards the interpretability of these models, understanding internal representations of LMMs remains largely a mystery. In this paper, we present a novel framework for the interpretation of LMMs. We propose a dictionary learning based approach, applied to the representation of tokens. The elements of the learned dictionary correspond to our proposed concepts. We show that these concepts are well semantically grounded in both vision and text. Thus we refer to these as ``multi-modal concepts''. We qualitatively and quantitatively evaluate the results of the learnt concepts. We show that the extracted multimodal concepts are useful to interpret representations of test samples. Finally, we evaluate the disentanglement between different concepts and the quality of grounding concepts visually and textually. Our code is publicly available at https://github.com/mshukor/xl-vlms
title A Concept-Based Explainability Framework for Large Multimodal Models
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.08074