SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mondal, Ishani, Li, Zongxia, Hou, Yufang, Natarajan, Anandhavelu, Garimella, Aparna, Boyd-Graber, Jordan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913548019433472
author Mondal, Ishani
Li, Zongxia
Hou, Yufang
Natarajan, Anandhavelu
Garimella, Aparna
Boyd-Graber, Jordan
author_facet Mondal, Ishani
Li, Zongxia
Hou, Yufang
Natarajan, Anandhavelu
Garimella, Aparna
Boyd-Graber, Jordan
contents Automating the creation of scientific diagrams from academic papers can significantly streamline the development of tutorials, presentations, and posters, thereby saving time and accelerating the process. Current text-to-image models struggle with generating accurate and visually appealing diagrams from long-context inputs. We propose SciDoc2Diagram, a task that extracts relevant information from scientific papers and generates diagrams, along with a benchmarking dataset, SciDoc2DiagramBench. We develop a multi-step pipeline SciDoc2Diagrammer that generates diagrams based on user intentions using intermediate code generation. We observed that initial diagram drafts were often incomplete or unfaithful to the source, leading us to develop SciDoc2Diagrammer-Multi-Aspect-Feedback (MAF), a refinement strategy that significantly enhances factual correctness and visual appeal and outperforms existing models on both automatic and human judgement.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19242
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
Mondal, Ishani
Li, Zongxia
Hou, Yufang
Natarajan, Anandhavelu
Garimella, Aparna
Boyd-Graber, Jordan
Computation and Language
Automating the creation of scientific diagrams from academic papers can significantly streamline the development of tutorials, presentations, and posters, thereby saving time and accelerating the process. Current text-to-image models struggle with generating accurate and visually appealing diagrams from long-context inputs. We propose SciDoc2Diagram, a task that extracts relevant information from scientific papers and generates diagrams, along with a benchmarking dataset, SciDoc2DiagramBench. We develop a multi-step pipeline SciDoc2Diagrammer that generates diagrams based on user intentions using intermediate code generation. We observed that initial diagram drafts were often incomplete or unfaithful to the source, leading us to develop SciDoc2Diagrammer-Multi-Aspect-Feedback (MAF), a refinement strategy that significantly enhances factual correctness and visual appeal and outperforms existing models on both automatic and human judgement.
title SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
topic Computation and Language
url https://arxiv.org/abs/2409.19242