Generating Animated Layouts as Structured Text Representations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shin, Yeonsang, Kim, Jihwan, Song, Yumin, Lee, Kyungseung, Chung, Hyunhee, Na, Taeyoung
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915270139838464
author Shin, Yeonsang
Kim, Jihwan
Song, Yumin
Lee, Kyungseung
Chung, Hyunhee
Na, Taeyoung
author_facet Shin, Yeonsang
Kim, Jihwan
Song, Yumin
Lee, Kyungseung
Chung, Hyunhee
Na, Taeyoung
contents Despite the remarkable progress in text-to-video models, achieving precise control over text elements and animated graphics remains a significant challenge, especially in applications such as video advertisements. To address this limitation, we introduce Animated Layout Generation, a novel approach to extend static graphic layouts with temporal dynamics. We propose a Structured Text Representation for fine-grained video control through hierarchical visual elements. To demonstrate the effectiveness of our approach, we present VAKER (Video Ad maKER), a text-to-video advertisement generation pipeline that combines a three-stage generation process with Unstructured Text Reasoning for seamless integration with LLMs. VAKER fully automates video advertisement generation by incorporating dynamic layout trajectories for objects and graphics across specific video frames. Through extensive evaluations, we demonstrate that VAKER significantly outperforms existing methods in generating video advertisements. Project Page: https://yeonsangshin.github.io/projects/Vaker
format Preprint
id arxiv_https___arxiv_org_abs_2505_00975
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generating Animated Layouts as Structured Text Representations
Shin, Yeonsang
Kim, Jihwan
Song, Yumin
Lee, Kyungseung
Chung, Hyunhee
Na, Taeyoung
Computer Vision and Pattern Recognition
Despite the remarkable progress in text-to-video models, achieving precise control over text elements and animated graphics remains a significant challenge, especially in applications such as video advertisements. To address this limitation, we introduce Animated Layout Generation, a novel approach to extend static graphic layouts with temporal dynamics. We propose a Structured Text Representation for fine-grained video control through hierarchical visual elements. To demonstrate the effectiveness of our approach, we present VAKER (Video Ad maKER), a text-to-video advertisement generation pipeline that combines a three-stage generation process with Unstructured Text Reasoning for seamless integration with LLMs. VAKER fully automates video advertisement generation by incorporating dynamic layout trajectories for objects and graphics across specific video frames. Through extensive evaluations, we demonstrate that VAKER significantly outperforms existing methods in generating video advertisements. Project Page: https://yeonsangshin.github.io/projects/Vaker
title Generating Animated Layouts as Structured Text Representations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.00975