T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Guo, Xuyang, Huo, Jiayan, Shi, Zhenmei, Song, Zhao, Zhang, Jiahao, Zhao, Jiale
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916726167306240
author Guo, Xuyang
Huo, Jiayan
Shi, Zhenmei
Song, Zhao
Zhang, Jiahao
Zhao, Jiale
author_facet Guo, Xuyang
Huo, Jiayan
Shi, Zhenmei
Song, Zhao
Zhang, Jiahao
Zhao, Jiale
contents Thanks to recent advancements in scalable deep architectures and large-scale pretraining, text-to-video generation has achieved unprecedented capabilities in producing high-fidelity, instruction-following content across a wide range of styles, enabling applications in advertising, entertainment, and education. However, these models' ability to render precise on-screen text, such as captions or mathematical formulas, remains largely untested, posing significant challenges for applications requiring exact textual accuracy. In this work, we introduce T2VTextBench, the first human-evaluation benchmark dedicated to evaluating on-screen text fidelity and temporal consistency in text-to-video models. Our suite of prompts integrates complex text strings with dynamic scene changes, testing each model's ability to maintain detailed instructions across frames. We evaluate ten state-of-the-art systems, ranging from open-source solutions to commercial offerings, and find that most struggle to generate legible, consistent text. These results highlight a critical gap in current video generators and provide a clear direction for future research aimed at enhancing textual manipulation in video synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04946
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
Guo, Xuyang
Huo, Jiayan
Shi, Zhenmei
Song, Zhao
Zhang, Jiahao
Zhao, Jiale
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Thanks to recent advancements in scalable deep architectures and large-scale pretraining, text-to-video generation has achieved unprecedented capabilities in producing high-fidelity, instruction-following content across a wide range of styles, enabling applications in advertising, entertainment, and education. However, these models' ability to render precise on-screen text, such as captions or mathematical formulas, remains largely untested, posing significant challenges for applications requiring exact textual accuracy. In this work, we introduce T2VTextBench, the first human-evaluation benchmark dedicated to evaluating on-screen text fidelity and temporal consistency in text-to-video models. Our suite of prompts integrates complex text strings with dynamic scene changes, testing each model's ability to maintain detailed instructions across frames. We evaluate ten state-of-the-art systems, ranging from open-source solutions to commercial offerings, and find that most struggle to generate legible, consistent text. These results highlight a critical gap in current video generators and provide a clear direction for future research aimed at enhancing textual manipulation in video synthesis.
title T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.04946