T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866916726167306240 |
|---|---|
| author | Guo, Xuyang Huo, Jiayan Shi, Zhenmei Song, Zhao Zhang, Jiahao Zhao, Jiale |
| author_facet | Guo, Xuyang Huo, Jiayan Shi, Zhenmei Song, Zhao Zhang, Jiahao Zhao, Jiale |
| contents | Thanks to recent advancements in scalable deep architectures and large-scale pretraining, text-to-video generation has achieved unprecedented capabilities in producing high-fidelity, instruction-following content across a wide range of styles, enabling applications in advertising, entertainment, and education. However, these models' ability to render precise on-screen text, such as captions or mathematical formulas, remains largely untested, posing significant challenges for applications requiring exact textual accuracy. In this work, we introduce T2VTextBench, the first human-evaluation benchmark dedicated to evaluating on-screen text fidelity and temporal consistency in text-to-video models. Our suite of prompts integrates complex text strings with dynamic scene changes, testing each model's ability to maintain detailed instructions across frames. We evaluate ten state-of-the-art systems, ranging from open-source solutions to commercial offerings, and find that most struggle to generate legible, consistent text. These results highlight a critical gap in current video generators and provide a clear direction for future research aimed at enhancing textual manipulation in video synthesis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_04946 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Guo, Xuyang Huo, Jiayan Shi, Zhenmei Song, Zhao Zhang, Jiahao Zhao, Jiale Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning Thanks to recent advancements in scalable deep architectures and large-scale pretraining, text-to-video generation has achieved unprecedented capabilities in producing high-fidelity, instruction-following content across a wide range of styles, enabling applications in advertising, entertainment, and education. However, these models' ability to render precise on-screen text, such as captions or mathematical formulas, remains largely untested, posing significant challenges for applications requiring exact textual accuracy. In this work, we introduce T2VTextBench, the first human-evaluation benchmark dedicated to evaluating on-screen text fidelity and temporal consistency in text-to-video models. Our suite of prompts integrates complex text strings with dynamic scene changes, testing each model's ability to maintain detailed instructions across frames. We evaluate ten state-of-the-art systems, ranging from open-source solutions to commercial offerings, and find that most struggle to generate legible, consistent text. These results highlight a critical gap in current video generators and provide a clear direction for future research aimed at enhancing textual manipulation in video synthesis. |
| title | T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2505.04946 |