Saved in:
Bibliographic Details
Main Authors: Lin, Zihao, Basu, Samyadeep, Beigi, Mohammad, Manjunatha, Varun, Rossi, Ryan A., Wang, Zichao, Zhou, Yufan, Balasubramanian, Sriram, Zarei, Arman, Rezaei, Keivan, Shen, Ying, Yao, Barry Menglong, Xu, Zhiyang, Liu, Qin, Zhang, Yuxiang, Sun, Yan, Liu, Shilong, Shen, Li, Li, Hongxuan, Feizi, Soheil, Huang, Lifu
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.17516
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929729890680832
author Lin, Zihao
Basu, Samyadeep
Beigi, Mohammad
Manjunatha, Varun
Rossi, Ryan A.
Wang, Zichao
Zhou, Yufan
Balasubramanian, Sriram
Zarei, Arman
Rezaei, Keivan
Shen, Ying
Yao, Barry Menglong
Xu, Zhiyang
Liu, Qin
Zhang, Yuxiang
Sun, Yan
Liu, Shilong
Shen, Li
Li, Hongxuan
Feizi, Soheil
Huang, Lifu
author_facet Lin, Zihao
Basu, Samyadeep
Beigi, Mohammad
Manjunatha, Varun
Rossi, Ryan A.
Wang, Zichao
Zhou, Yufan
Balasubramanian, Sriram
Zarei, Arman
Rezaei, Keivan
Shen, Ying
Yao, Barry Menglong
Xu, Zhiyang
Liu, Qin
Zhang, Yuxiang
Sun, Yan
Liu, Shilong
Shen, Li
Li, Hongxuan
Feizi, Soheil
Huang, Lifu
contents The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for better control. While significant progress has been made in interpreting Large Language Models (LLMs), multimodal foundation models (MMFMs) - such as contrastive vision-language models, generative vision-language models, and text-to-image models - pose unique interpretability challenges beyond unimodal frameworks. Despite initial studies, a substantial gap remains between the interpretability of LLMs and MMFMs. This survey explores two key aspects: (1) the adaptation of LLM interpretability methods to multimodal models and (2) understanding the mechanistic differences between unimodal language models and crossmodal systems. By systematically reviewing current MMFM analysis techniques, we propose a structured taxonomy of interpretability methods, compare insights across unimodal and multimodal architectures, and highlight critical research gaps.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17516
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
Lin, Zihao
Basu, Samyadeep
Beigi, Mohammad
Manjunatha, Varun
Rossi, Ryan A.
Wang, Zichao
Zhou, Yufan
Balasubramanian, Sriram
Zarei, Arman
Rezaei, Keivan
Shen, Ying
Yao, Barry Menglong
Xu, Zhiyang
Liu, Qin
Zhang, Yuxiang
Sun, Yan
Liu, Shilong
Shen, Li
Li, Hongxuan
Feizi, Soheil
Huang, Lifu
Machine Learning
Artificial Intelligence
The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for better control. While significant progress has been made in interpreting Large Language Models (LLMs), multimodal foundation models (MMFMs) - such as contrastive vision-language models, generative vision-language models, and text-to-image models - pose unique interpretability challenges beyond unimodal frameworks. Despite initial studies, a substantial gap remains between the interpretability of LLMs and MMFMs. This survey explores two key aspects: (1) the adaptation of LLM interpretability methods to multimodal models and (2) understanding the mechanistic differences between unimodal language models and crossmodal systems. By systematically reviewing current MMFM analysis techniques, we propose a structured taxonomy of interpretability methods, compare insights across unimodal and multimodal architectures, and highlight critical research gaps.
title A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.17516