PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jaisankar, Vijay, Bandyopadhyay, Sambaran, Vyas, Kalp, Chaitanya, Varre, Somasundaram, Shwetha
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913370914947072
author Jaisankar, Vijay
Bandyopadhyay, Sambaran
Vyas, Kalp
Chaitanya, Varre
Somasundaram, Shwetha
author_facet Jaisankar, Vijay
Bandyopadhyay, Sambaran
Vyas, Kalp
Chaitanya, Varre
Somasundaram, Shwetha
contents A poster from a long input document can be considered as a one-page easy-to-read multimodal (text and images) summary presented on a nice template with good design elements. Automatic transformation of a long document into a poster is a very less studied but challenging task. It involves content summarization of the input document followed by template generation and harmonization. In this work, we propose a novel deep submodular function which can be trained on ground truth summaries to extract multimodal content from the document and explicitly ensures good coverage, diversity and alignment of text and images. Then, we use an LLM based paraphraser and propose to generate a template with various design aspects conditioned on the input content. We show the merits of our approach through extensive automated and human evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20213
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization
Jaisankar, Vijay
Bandyopadhyay, Sambaran
Vyas, Kalp
Chaitanya, Varre
Somasundaram, Shwetha
Artificial Intelligence
Computation and Language
Machine Learning
A poster from a long input document can be considered as a one-page easy-to-read multimodal (text and images) summary presented on a nice template with good design elements. Automatic transformation of a long document into a poster is a very less studied but challenging task. It involves content summarization of the input document followed by template generation and harmonization. In this work, we propose a novel deep submodular function which can be trained on ground truth summaries to extract multimodal content from the document and explicitly ensures good coverage, diversity and alignment of text and images. Then, we use an LLM based paraphraser and propose to generate a template with various design aspects conditioned on the input content. We show the merits of our approach through extensive automated and human evaluations.
title PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2405.20213