BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Yuyang, Xiao, Shishi, Wu, Keming, Liao, Qisheng, Chen, Bohan, Lin, Kevin, Huang, Danqing, Li, Ji, Yuan, Yuhui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918079811813376
author Peng, Yuyang
Xiao, Shishi
Wu, Keming
Liao, Qisheng
Chen, Bohan
Lin, Kevin
Huang, Danqing
Li, Ji
Yuan, Yuhui
author_facet Peng, Yuyang
Xiao, Shishi
Wu, Keming
Liao, Qisheng
Chen, Bohan
Lin, Kevin
Huang, Danqing
Li, Ji
Yuan, Yuhui
contents Recently, state-of-the-art text-to-image generation models, such as Flux and Ideogram 2.0, have made significant progress in sentence-level visual text rendering. In this paper, we focus on the more challenging scenarios of article-level visual text rendering and address a novel task of generating high-quality business content, including infographics and slides, based on user provided article-level descriptive prompts and ultra-dense layouts. The fundamental challenges are twofold: significantly longer context lengths and the scarcity of high-quality business content data. In contrast to most previous works that focus on a limited number of sub-regions and sentence-level prompts, ensuring precise adherence to ultra-dense layouts with tens or even hundreds of sub-regions in business content is far more challenging. We make two key technical contributions: (i) the construction of scalable, high-quality business content dataset, i.e., Infographics-650K, equipped with ultra-dense layouts and prompts by implementing a layer-wise retrieval-augmented infographic generation scheme; and (ii) a layout-guided cross attention scheme, which injects tens of region-wise prompts into a set of cropped region latent space according to the ultra-dense layouts, and refine each sub-regions flexibly during inference using a layout conditional CFG. We demonstrate the strong results of our system compared to previous SOTA systems such as Flux and SD3 on our BizEval prompt set. Additionally, we conduct thorough ablation experiments to verify the effectiveness of each component. We hope our constructed Infographics-650K and BizEval can encourage the broader community to advance the progress of business content generation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20672
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation
Peng, Yuyang
Xiao, Shishi
Wu, Keming
Liao, Qisheng
Chen, Bohan
Lin, Kevin
Huang, Danqing
Li, Ji
Yuan, Yuhui
Computer Vision and Pattern Recognition
Recently, state-of-the-art text-to-image generation models, such as Flux and Ideogram 2.0, have made significant progress in sentence-level visual text rendering. In this paper, we focus on the more challenging scenarios of article-level visual text rendering and address a novel task of generating high-quality business content, including infographics and slides, based on user provided article-level descriptive prompts and ultra-dense layouts. The fundamental challenges are twofold: significantly longer context lengths and the scarcity of high-quality business content data. In contrast to most previous works that focus on a limited number of sub-regions and sentence-level prompts, ensuring precise adherence to ultra-dense layouts with tens or even hundreds of sub-regions in business content is far more challenging. We make two key technical contributions: (i) the construction of scalable, high-quality business content dataset, i.e., Infographics-650K, equipped with ultra-dense layouts and prompts by implementing a layer-wise retrieval-augmented infographic generation scheme; and (ii) a layout-guided cross attention scheme, which injects tens of region-wise prompts into a set of cropped region latent space according to the ultra-dense layouts, and refine each sub-regions flexibly during inference using a layout conditional CFG. We demonstrate the strong results of our system compared to previous SOTA systems such as Flux and SD3 on our BizEval prompt set. Additionally, we conduct thorough ablation experiments to verify the effectiveness of each component. We hope our constructed Infographics-650K and BizEval can encourage the broader community to advance the progress of business content generation.
title BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.20672