ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Bowen, Yuan, Yuan, Bai, Xinyi, Hao, Zhuoqun, Yin, Alyson, Hu, Yaojie, Liao, Wenyu, Ungar, Lyle, Taylor, Camillo J.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908613259296768
author Jiang, Bowen
Yuan, Yuan
Bai, Xinyi
Hao, Zhuoqun
Yin, Alyson
Hu, Yaojie
Liao, Wenyu
Ungar, Lyle
Taylor, Camillo J.
author_facet Jiang, Bowen
Yuan, Yuan
Bai, Xinyi
Hao, Zhuoqun
Yin, Alyson
Hu, Yaojie
Liao, Wenyu
Ungar, Lyle
Taylor, Camillo J.
contents This work demonstrates that diffusion models can achieve font-controllable multilingual text rendering using just raw images without font label annotations.Visual text rendering remains a significant challenge. While recent methods condition diffusion on glyphs, it is impossible to retrieve exact font annotations from large-scale, real-world datasets, which prevents user-specified font control. To address this, we propose a data-driven solution that integrates the conditional diffusion model with a text segmentation model, utilizing segmentation masks to capture and represent fonts in pixel space in a self-supervised manner, thereby eliminating the need for any ground-truth labels and enabling users to customize text rendering with any multilingual font of their choice. The experiment provides a proof of concept of our algorithm in zero-shot text and font editing across diverse fonts and languages, providing valuable insights for the community and industry toward achieving generalized visual text rendering. Code is available at github.com/bowen-upenn/ControlText.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10999
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations
Jiang, Bowen
Yuan, Yuan
Bai, Xinyi
Hao, Zhuoqun
Yin, Alyson
Hu, Yaojie
Liao, Wenyu
Ungar, Lyle
Taylor, Camillo J.
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
This work demonstrates that diffusion models can achieve font-controllable multilingual text rendering using just raw images without font label annotations.Visual text rendering remains a significant challenge. While recent methods condition diffusion on glyphs, it is impossible to retrieve exact font annotations from large-scale, real-world datasets, which prevents user-specified font control. To address this, we propose a data-driven solution that integrates the conditional diffusion model with a text segmentation model, utilizing segmentation masks to capture and represent fonts in pixel space in a self-supervised manner, thereby eliminating the need for any ground-truth labels and enabling users to customize text rendering with any multilingual font of their choice. The experiment provides a proof of concept of our algorithm in zero-shot text and font editing across diverse fonts and languages, providing valuable insights for the community and industry toward achieving generalized visual text rendering. Code is available at github.com/bowen-upenn/ControlText.
title ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
url https://arxiv.org/abs/2502.10999