Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908895171051520 |
|---|---|
| author | Chen, Ruoyu Guo, Xiaoqing Liu, Kangwei Liang, Siyuan Liu, Shiming Zhang, Qunli Wang, Laiyuan Zhang, Hua Cao, Xiaochun |
| author_facet | Chen, Ruoyu Guo, Xiaoqing Liu, Kangwei Liang, Siyuan Liu, Shiming Zhang, Qunli Wang, Laiyuan Zhang, Hua Cao, Xiaochun |
| contents | Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated tokens depend on visual modalities remains poorly understood, limiting interpretability and reliability. In this work, we present EAGLE, a lightweight black-box framework for explaining autoregressive token generation in MLLMs. EAGLE attributes any selected tokens to compact perceptual regions while quantifying the relative influence of language priors and perceptual evidence. The framework introduces an objective function that unifies sufficiency (insight score) and indispensability (necessity score), optimized via greedy search over sparsified image regions for faithful and efficient attribution. Beyond spatial attribution, EAGLE performs modality-aware analysis that disentangles what tokens rely on, providing fine-grained interpretability of model decisions. Extensive experiments across open-source MLLMs show that EAGLE consistently outperforms existing methods in faithfulness, localization, and hallucination diagnosis, while requiring substantially less GPU memory. These results highlight its effectiveness and practicality for advancing the interpretability of MLLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_22496 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation Chen, Ruoyu Guo, Xiaoqing Liu, Kangwei Liang, Siyuan Liu, Shiming Zhang, Qunli Wang, Laiyuan Zhang, Hua Cao, Xiaochun Computer Vision and Pattern Recognition Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated tokens depend on visual modalities remains poorly understood, limiting interpretability and reliability. In this work, we present EAGLE, a lightweight black-box framework for explaining autoregressive token generation in MLLMs. EAGLE attributes any selected tokens to compact perceptual regions while quantifying the relative influence of language priors and perceptual evidence. The framework introduces an objective function that unifies sufficiency (insight score) and indispensability (necessity score), optimized via greedy search over sparsified image regions for faithful and efficient attribution. Beyond spatial attribution, EAGLE performs modality-aware analysis that disentangles what tokens rely on, providing fine-grained interpretability of model decisions. Extensive experiments across open-source MLLMs show that EAGLE consistently outperforms existing methods in faithfulness, localization, and hallucination diagnosis, while requiring substantially less GPU memory. These results highlight its effectiveness and practicality for advancing the interpretability of MLLMs. |
| title | Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.22496 |