We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Choi, Minkyu, Sharan, S P, Goel, Harsh, Shah, Sahil, Chinchali, Sandeep
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918425864962048
author Choi, Minkyu
Sharan, S P
Goel, Harsh
Shah, Sahil
Chinchali, Sandeep
author_facet Choi, Minkyu
Sharan, S P
Goel, Harsh
Shah, Sahil
Chinchali, Sandeep
contents Current text-to-video (T2V) generation models are increasingly popular due to their ability to produce coherent videos from textual prompts. However, these models often struggle to generate semantically and temporally consistent videos when dealing with longer, more complex prompts involving multiple objects or sequential events. Additionally, the high computational costs associated with training or fine-tuning make direct improvements impractical. To overcome these limitations, we introduce NeuS-E, a novel zero-training video refinement pipeline that leverages neuro-symbolic feedback to automatically enhance video generation, achieving superior alignment with the prompts. Our approach first derives the neuro-symbolic feedback by analyzing a formal video representation and pinpoints semantically inconsistent events, objects, and their corresponding frames. This feedback then guides targeted edits to the original video. Extensive empirical evaluations on both open-source and proprietary T2V models demonstrate that NeuS-E significantly enhances temporal and logical alignment across diverse prompts by almost 40%
format Preprint
id arxiv_https___arxiv_org_abs_2504_17180
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback
Choi, Minkyu
Sharan, S P
Goel, Harsh
Shah, Sahil
Chinchali, Sandeep
Computer Vision and Pattern Recognition
Artificial Intelligence
Current text-to-video (T2V) generation models are increasingly popular due to their ability to produce coherent videos from textual prompts. However, these models often struggle to generate semantically and temporally consistent videos when dealing with longer, more complex prompts involving multiple objects or sequential events. Additionally, the high computational costs associated with training or fine-tuning make direct improvements impractical. To overcome these limitations, we introduce NeuS-E, a novel zero-training video refinement pipeline that leverages neuro-symbolic feedback to automatically enhance video generation, achieving superior alignment with the prompts. Our approach first derives the neuro-symbolic feedback by analyzing a formal video representation and pinpoints semantically inconsistent events, objects, and their corresponding frames. This feedback then guides targeted edits to the original video. Extensive empirical evaluations on both open-source and proprietary T2V models demonstrate that NeuS-E significantly enhances temporal and logical alignment across diverse prompts by almost 40%
title We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.17180