Large Multi-modal Models Can Interpret Features in Large Multi-modal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kaichen, Shen, Yifei, Li, Bo, Liu, Ziwei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915499831459840
author Zhang, Kaichen
Shen, Yifei
Li, Bo
Liu, Ziwei
author_facet Zhang, Kaichen
Shen, Yifei
Li, Bo
Liu, Ziwei
contents Recent advances in Large Multimodal Models (LMMs) lead to significant breakthroughs in both academia and industry. One question that arises is how we, as humans, can understand their internal neural representations. This paper takes an initial step towards addressing this question by presenting a versatile framework to identify and interpret the semantics within LMMs. Specifically, 1) we first apply a Sparse Autoencoder(SAE) to disentangle the representations into human understandable features. 2) We then present an automatic interpretation framework to interpreted the open-semantic features learned in SAE by the LMMs themselves. We employ this framework to analyze the LLaVA-NeXT-8B model using the LLaVA-OV-72B model, demonstrating that these features can effectively steer the model's behavior. Our results contribute to a deeper understanding of why LMMs excel in specific tasks, including EQ tests, and illuminate the nature of their mistakes along with potential strategies for their rectification. These findings offer new insights into the internal mechanisms of LMMs and suggest parallels with the cognitive processes of the human brain.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14982
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
Zhang, Kaichen
Shen, Yifei
Li, Bo
Liu, Ziwei
Computer Vision and Pattern Recognition
Computation and Language
Recent advances in Large Multimodal Models (LMMs) lead to significant breakthroughs in both academia and industry. One question that arises is how we, as humans, can understand their internal neural representations. This paper takes an initial step towards addressing this question by presenting a versatile framework to identify and interpret the semantics within LMMs. Specifically, 1) we first apply a Sparse Autoencoder(SAE) to disentangle the representations into human understandable features. 2) We then present an automatic interpretation framework to interpreted the open-semantic features learned in SAE by the LMMs themselves. We employ this framework to analyze the LLaVA-NeXT-8B model using the LLaVA-OV-72B model, demonstrating that these features can effectively steer the model's behavior. Our results contribute to a deeper understanding of why LMMs excel in specific tasks, including EQ tests, and illuminate the nature of their mistakes along with potential strategies for their rectification. These findings offer new insights into the internal mechanisms of LMMs and suggest parallels with the cognitive processes of the human brain.
title Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2411.14982