PTTA: A Pure Text-to-Animation Framework for High-Quality Creation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Ruiqi, Cai, Kaitong, Fan, Yijia, Wang, Keze
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912780251037696
author Chen, Ruiqi
Cai, Kaitong
Fan, Yijia
Wang, Keze
author_facet Chen, Ruiqi
Cai, Kaitong
Fan, Yijia
Wang, Keze
contents Traditional animation production involves complex pipelines and significant manual labor cost. While recent video generation models such as Sora, Kling, and CogVideoX achieve impressive results on natural video synthesis, they exhibit notable limitations when applied to animation generation. Recent efforts, such as AniSora, demonstrate promising performance by fine-tuning image-to-video models for animation styles, yet analogous exploration in the text-to-video setting remains limited. In this work, we present PTTA, a pure text-to-animation framework for high-quality animation creation. We first construct a small-scale but high-quality paired dataset of animation videos and textual descriptions. Building upon the pretrained text-to-video model HunyuanVideo, we perform fine-tuning to adapt it to animation-style generation. Extensive visual evaluations across multiple dimensions show that the proposed approach consistently outperforms comparable baselines in animation video synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18614
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PTTA: A Pure Text-to-Animation Framework for High-Quality Creation
Chen, Ruiqi
Cai, Kaitong
Fan, Yijia
Wang, Keze
Computer Vision and Pattern Recognition
Artificial Intelligence
Traditional animation production involves complex pipelines and significant manual labor cost. While recent video generation models such as Sora, Kling, and CogVideoX achieve impressive results on natural video synthesis, they exhibit notable limitations when applied to animation generation. Recent efforts, such as AniSora, demonstrate promising performance by fine-tuning image-to-video models for animation styles, yet analogous exploration in the text-to-video setting remains limited. In this work, we present PTTA, a pure text-to-animation framework for high-quality animation creation. We first construct a small-scale but high-quality paired dataset of animation videos and textual descriptions. Building upon the pretrained text-to-video model HunyuanVideo, we perform fine-tuning to adapt it to animation-style generation. Extensive visual evaluations across multiple dimensions show that the proposed approach consistently outperforms comparable baselines in animation video synthesis.
title PTTA: A Pure Text-to-Animation Framework for High-Quality Creation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.18614