Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gigant, Théo, Guinaudeau, Camille, Dufaux, Frédéric
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917984983842816
author Gigant, Théo
Guinaudeau, Camille
Dufaux, Frédéric
author_facet Gigant, Théo
Guinaudeau, Camille
Dufaux, Frédéric
contents Vision-Language Models (VLMs) can process visual and textual information in multiple formats: texts, images, interleaved texts and images, or even hour-long videos. In this work, we conduct fine-grained quantitative and qualitative analyses of automatic summarization of multimodal presentations using VLMs with various representations as input. From these experiments, we suggest cost-effective strategies for generating summaries from text-heavy multimodal documents under different input-length budgets using VLMs. We show that slides extracted from the video stream can be beneficially used as input against the raw video, and that a structured representation from interleaved slides and transcript provides the best performance. Finally, we reflect and comment on the nature of cross-modal interactions in multimodal presentations and share suggestions to improve the capabilities of VLMs to understand documents of this nature.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
Gigant, Théo
Guinaudeau, Camille
Dufaux, Frédéric
Computer Vision and Pattern Recognition
Computation and Language
Vision-Language Models (VLMs) can process visual and textual information in multiple formats: texts, images, interleaved texts and images, or even hour-long videos. In this work, we conduct fine-grained quantitative and qualitative analyses of automatic summarization of multimodal presentations using VLMs with various representations as input. From these experiments, we suggest cost-effective strategies for generating summaries from text-heavy multimodal documents under different input-length budgets using VLMs. We show that slides extracted from the video stream can be beneficially used as input against the raw video, and that a structured representation from interleaved slides and transcript provides the best performance. Finally, we reflect and comment on the nature of cross-modal interactions in multimodal presentations and share suggestions to improve the capabilities of VLMs to understand documents of this nature.
title Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2504.10049