GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Zexuan, Jin, Jiarui, Ma, Yue, Wang, Shijian, Hu, Jiahui, Jiao, Wenxiang, Lu, Yuan, Zhang, Linfeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911509271019520
author Yan, Zexuan
Jin, Jiarui
Ma, Yue
Wang, Shijian
Hu, Jiahui
Jiao, Wenxiang
Lu, Yuan
Zhang, Linfeng
author_facet Yan, Zexuan
Jin, Jiarui
Ma, Yue
Wang, Shijian
Hu, Jiahui
Jiao, Wenxiang
Lu, Yuan
Zhang, Linfeng
contents Despite recent advances in generative models driving significant progress in text rendering, accurately generating complex text and mathematical formulas remains a formidable challenge. This difficulty primarily stems from the limited instruction-following capabilities of current models when encountering out-of-distribution prompts. To address this, we introduce GlyphBanana, alongside a corresponding benchmark specifically designed for rendering complex characters and formulas. GlyphBanana employs an agentic workflow that integrates auxiliary tools to inject glyph templates into both the latent space and attention maps, facilitating the iterative refinement of generated images. Notably, our training-free approach can be seamlessly applied to various Text-to-Image (T2I) models, achieving superior precision compared to existing baselines. Extensive experiments demonstrate the effectiveness of our proposed workflow. Associated code is publicly available at https://github.com/yuriYanZeXuan/GlyphBanana.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12155
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows
Yan, Zexuan
Jin, Jiarui
Ma, Yue
Wang, Shijian
Hu, Jiahui
Jiao, Wenxiang
Lu, Yuan
Zhang, Linfeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Despite recent advances in generative models driving significant progress in text rendering, accurately generating complex text and mathematical formulas remains a formidable challenge. This difficulty primarily stems from the limited instruction-following capabilities of current models when encountering out-of-distribution prompts. To address this, we introduce GlyphBanana, alongside a corresponding benchmark specifically designed for rendering complex characters and formulas. GlyphBanana employs an agentic workflow that integrates auxiliary tools to inject glyph templates into both the latent space and attention maps, facilitating the iterative refinement of generated images. Notably, our training-free approach can be seamlessly applied to various Text-to-Image (T2I) models, achieving superior precision compared to existing baselines. Extensive experiments demonstrate the effectiveness of our proposed workflow. Associated code is publicly available at https://github.com/yuriYanZeXuan/GlyphBanana.
title GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.12155