Planning with Sketch-Guided Verification for Physics-Aware Video Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Huang, Yidong, Wang, Zun, Lin, Han, Kim, Dong-Ki, Omidshafiei, Shayegan, Yoon, Jaehong, Zhang, Yue, Bansal, Mohit
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908669048782848
author Huang, Yidong
Wang, Zun
Lin, Han
Kim, Dong-Ki
Omidshafiei, Shayegan
Yoon, Jaehong
Zhang, Yue
Bansal, Mohit
author_facet Huang, Yidong
Wang, Zun
Lin, Han
Kim, Dong-Ki
Omidshafiei, Shayegan
Yoon, Jaehong
Zhang, Yue
Bansal, Mohit
contents Recent video generation approaches increasingly rely on planning intermediate control signals such as object trajectories to improve temporal coherence and motion fidelity. However, these methods mostly employ single-shot plans that are typically limited to simple motions, or iterative refinement which requires multiple calls to the video generator, incuring high computational cost. To overcome these limitations, we propose SketchVerify, a training-free, sketch-verification-based planning framework that improves motion planning quality with more dynamically coherent trajectories (i.e., physically plausible and instruction-consistent motions) prior to full video generation by introducing a test-time sampling and verification loop. Given a prompt and a reference image, our method predicts multiple candidate motion plans and ranks them using a vision-language verifier that jointly evaluates semantic alignment with the instruction and physical plausibility. To efficiently score candidate motion plans, we render each trajectory as a lightweight video sketch by compositing objects over a static background, which bypasses the need for expensive, repeated diffusion-based synthesis while achieving comparable performance. We iteratively refine the motion plan until a satisfactory one is identified, which is then passed to the trajectory-conditioned generator for final synthesis. Experiments on WorldModelBench and PhyWorldBench demonstrate that our method significantly improves motion quality, physical realism, and long-term consistency compared to competitive baselines while being substantially more efficient. Our ablation study further shows that scaling up the number of trajectory candidates consistently enhances overall performance.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17450
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Planning with Sketch-Guided Verification for Physics-Aware Video Generation
Huang, Yidong
Wang, Zun
Lin, Han
Kim, Dong-Ki
Omidshafiei, Shayegan
Yoon, Jaehong
Zhang, Yue
Bansal, Mohit
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Recent video generation approaches increasingly rely on planning intermediate control signals such as object trajectories to improve temporal coherence and motion fidelity. However, these methods mostly employ single-shot plans that are typically limited to simple motions, or iterative refinement which requires multiple calls to the video generator, incuring high computational cost. To overcome these limitations, we propose SketchVerify, a training-free, sketch-verification-based planning framework that improves motion planning quality with more dynamically coherent trajectories (i.e., physically plausible and instruction-consistent motions) prior to full video generation by introducing a test-time sampling and verification loop. Given a prompt and a reference image, our method predicts multiple candidate motion plans and ranks them using a vision-language verifier that jointly evaluates semantic alignment with the instruction and physical plausibility. To efficiently score candidate motion plans, we render each trajectory as a lightweight video sketch by compositing objects over a static background, which bypasses the need for expensive, repeated diffusion-based synthesis while achieving comparable performance. We iteratively refine the motion plan until a satisfactory one is identified, which is then passed to the trajectory-conditioned generator for final synthesis. Experiments on WorldModelBench and PhyWorldBench demonstrate that our method significantly improves motion quality, physical realism, and long-term consistency compared to competitive baselines while being substantially more efficient. Our ablation study further shows that scaling up the number of trajectory candidates consistently enhances overall performance.
title Planning with Sketch-Guided Verification for Physics-Aware Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2511.17450