PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913370914947072 |
|---|---|
| author | Jaisankar, Vijay Bandyopadhyay, Sambaran Vyas, Kalp Chaitanya, Varre Somasundaram, Shwetha |
| author_facet | Jaisankar, Vijay Bandyopadhyay, Sambaran Vyas, Kalp Chaitanya, Varre Somasundaram, Shwetha |
| contents | A poster from a long input document can be considered as a one-page easy-to-read multimodal (text and images) summary presented on a nice template with good design elements. Automatic transformation of a long document into a poster is a very less studied but challenging task. It involves content summarization of the input document followed by template generation and harmonization. In this work, we propose a novel deep submodular function which can be trained on ground truth summaries to extract multimodal content from the document and explicitly ensures good coverage, diversity and alignment of text and images. Then, we use an LLM based paraphraser and propose to generate a template with various design aspects conditioned on the input content. We show the merits of our approach through extensive automated and human evaluations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_20213 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization Jaisankar, Vijay Bandyopadhyay, Sambaran Vyas, Kalp Chaitanya, Varre Somasundaram, Shwetha Artificial Intelligence Computation and Language Machine Learning A poster from a long input document can be considered as a one-page easy-to-read multimodal (text and images) summary presented on a nice template with good design elements. Automatic transformation of a long document into a poster is a very less studied but challenging task. It involves content summarization of the input document followed by template generation and harmonization. In this work, we propose a novel deep submodular function which can be trained on ground truth summaries to extract multimodal content from the document and explicitly ensures good coverage, diversity and alignment of text and images. Then, we use an LLM based paraphraser and propose to generate a template with various design aspects conditioned on the input content. We show the merits of our approach through extensive automated and human evaluations. |
| title | PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization |
| topic | Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2405.20213 |