Multimodal QUD: Inquisitive Questions from Scientific Figures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yating, Rudman, William, Govindarajan, Venkata S, Dimakis, Alexandros G., Li, Junyi Jessy
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915959419174912
author Wu, Yating
Rudman, William
Govindarajan, Venkata S
Dimakis, Alexandros G.
Li, Junyi Jessy
author_facet Wu, Yating
Rudman, William
Govindarajan, Venkata S
Dimakis, Alexandros G.
Li, Junyi Jessy
contents Asking inquisitive questions while reading, and looking for their answers, is an important part in human discourse comprehension, curiosity, and creative ideation, and prior work has investigated this in text-only scenarios. However, in scientific or research papers, many of the critical takeaways are conveyed through both figures and the text that analyzes them. While scientific visualizations have been used to evaluate Vision-Language Models (VLMs) capabilities, current benchmarks are limited to questions that focus simply on extracting information from them. Such questions only require lower-level reasoning, do not take into account the context in which a figure appears, and do not reflect the communicative goals the authors wish to achieve. We generate inquisitive questions that reach the depth of questions humans generate when engaging with scientific papers, conditioned on both the figure and the paper's context, and require reasoning across both modalities. To do so, we extend the linguistic theory of Questions Under Discussion (QUD) from being text-only to multimodal, where implicit questions are raised and resolved as discourse progresses. We present MQUD, a dataset of research papers in which such questions are made explicit and annotated by the original authors. We show that fine-tuning a VLM on MQUD shifts the model from generating generic low-level visual questions to content-specific grounding that requires a high-level of multimodal reasoning, yielding higher-quality, more visually grounded multimodal QUD generation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23733
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multimodal QUD: Inquisitive Questions from Scientific Figures
Wu, Yating
Rudman, William
Govindarajan, Venkata S
Dimakis, Alexandros G.
Li, Junyi Jessy
Computation and Language
Asking inquisitive questions while reading, and looking for their answers, is an important part in human discourse comprehension, curiosity, and creative ideation, and prior work has investigated this in text-only scenarios. However, in scientific or research papers, many of the critical takeaways are conveyed through both figures and the text that analyzes them. While scientific visualizations have been used to evaluate Vision-Language Models (VLMs) capabilities, current benchmarks are limited to questions that focus simply on extracting information from them. Such questions only require lower-level reasoning, do not take into account the context in which a figure appears, and do not reflect the communicative goals the authors wish to achieve. We generate inquisitive questions that reach the depth of questions humans generate when engaging with scientific papers, conditioned on both the figure and the paper's context, and require reasoning across both modalities. To do so, we extend the linguistic theory of Questions Under Discussion (QUD) from being text-only to multimodal, where implicit questions are raised and resolved as discourse progresses. We present MQUD, a dataset of research papers in which such questions are made explicit and annotated by the original authors. We show that fine-tuning a VLM on MQUD shifts the model from generating generic low-level visual questions to content-specific grounding that requires a high-level of multimodal reasoning, yielding higher-quality, more visually grounded multimodal QUD generation.
title Multimodal QUD: Inquisitive Questions from Scientific Figures
topic Computation and Language
url https://arxiv.org/abs/2604.23733