TikZero: Zero-Shot Text-Guided Graphics Program Synthesis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Belouadi, Jonas, Ilg, Eddy, Keuper, Margret, Tanaka, Hideki, Utiyama, Masao, Dabre, Raj, Eger, Steffen, Ponzetto, Simone Paolo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913989977440256
author Belouadi, Jonas
Ilg, Eddy
Keuper, Margret
Tanaka, Hideki
Utiyama, Masao
Dabre, Raj
Eger, Steffen
Ponzetto, Simone Paolo
author_facet Belouadi, Jonas
Ilg, Eddy
Keuper, Margret
Tanaka, Hideki
Utiyama, Masao
Dabre, Raj
Eger, Steffen
Ponzetto, Simone Paolo
contents Automatically synthesizing figures from text captions is a compelling capability. However, achieving high geometric precision and editability requires representing figures as graphics programs in languages like TikZ, and aligned training data (i.e., graphics programs with captions) remains scarce. Meanwhile, large amounts of unaligned graphics programs and captioned raster images are more readily available. We reconcile these disparate data sources by presenting TikZero, which decouples graphics program generation from text understanding by using image representations as an intermediary bridge. It enables independent training on graphics programs and captioned images and allows for zero-shot text-guided graphics program synthesis during inference. We show that our method substantially outperforms baselines that can only operate with caption-aligned graphics programs. Furthermore, when leveraging caption-aligned graphics programs as a complementary training signal, TikZero matches or exceeds the performance of much larger models, including commercial systems like GPT-4o. Our code, datasets, and select models are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11509
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TikZero: Zero-Shot Text-Guided Graphics Program Synthesis
Belouadi, Jonas
Ilg, Eddy
Keuper, Margret
Tanaka, Hideki
Utiyama, Masao
Dabre, Raj
Eger, Steffen
Ponzetto, Simone Paolo
Computation and Language
Computer Vision and Pattern Recognition
Automatically synthesizing figures from text captions is a compelling capability. However, achieving high geometric precision and editability requires representing figures as graphics programs in languages like TikZ, and aligned training data (i.e., graphics programs with captions) remains scarce. Meanwhile, large amounts of unaligned graphics programs and captioned raster images are more readily available. We reconcile these disparate data sources by presenting TikZero, which decouples graphics program generation from text understanding by using image representations as an intermediary bridge. It enables independent training on graphics programs and captioned images and allows for zero-shot text-guided graphics program synthesis during inference. We show that our method substantially outperforms baselines that can only operate with caption-aligned graphics programs. Furthermore, when leveraging caption-aligned graphics programs as a complementary training signal, TikZero matches or exceeds the performance of much larger models, including commercial systems like GPT-4o. Our code, datasets, and select models are publicly available.
title TikZero: Zero-Shot Text-Guided Graphics Program Synthesis
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.11509