Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Wang, Jia, Liyu, Hu, Wentao, Pan, Kaihang, Yue, Zhongqi, Zhao, Wei, Chen, Jingyuan, Wu, Fei, Zhang, Hanwang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908332284968960
author Lin, Wang
Jia, Liyu
Hu, Wentao
Pan, Kaihang
Yue, Zhongqi
Zhao, Wei
Chen, Jingyuan
Wu, Fei
Zhang, Hanwang
author_facet Lin, Wang
Jia, Liyu
Hu, Wentao
Pan, Kaihang
Yue, Zhongqi
Zhao, Wei
Chen, Jingyuan
Wu, Fei
Zhang, Hanwang
contents Despite recent progress in video generation, producing videos that adhere to physical laws remains a significant challenge. Traditional diffusion-based methods struggle to extrapolate to unseen physical conditions (eg, velocity) due to their reliance on data-driven approximations. To address this, we propose to integrate symbolic reasoning and reinforcement learning to enforce physical consistency in video generation. We first introduce the Diffusion Timestep Tokenizer (DDT), which learns discrete, recursive visual tokens by recovering visual attributes lost during the diffusion process. The recursive visual tokens enable symbolic reasoning by a large language model. Based on it, we propose the Phys-AR framework, which consists of two stages: The first stage uses supervised fine-tuning to transfer symbolic knowledge, while the second stage applies reinforcement learning to optimize the model's reasoning abilities through reward functions based on physical conditions. Our approach allows the model to dynamically adjust and improve the physical properties of generated videos, ensuring adherence to physical laws. Experimental results demonstrate that PhysAR can generate videos that are physically consistent.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15932
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
Lin, Wang
Jia, Liyu
Hu, Wentao
Pan, Kaihang
Yue, Zhongqi
Zhao, Wei
Chen, Jingyuan
Wu, Fei
Zhang, Hanwang
Computer Vision and Pattern Recognition
Despite recent progress in video generation, producing videos that adhere to physical laws remains a significant challenge. Traditional diffusion-based methods struggle to extrapolate to unseen physical conditions (eg, velocity) due to their reliance on data-driven approximations. To address this, we propose to integrate symbolic reasoning and reinforcement learning to enforce physical consistency in video generation. We first introduce the Diffusion Timestep Tokenizer (DDT), which learns discrete, recursive visual tokens by recovering visual attributes lost during the diffusion process. The recursive visual tokens enable symbolic reasoning by a large language model. Based on it, we propose the Phys-AR framework, which consists of two stages: The first stage uses supervised fine-tuning to transfer symbolic knowledge, while the second stage applies reinforcement learning to optimize the model's reasoning abilities through reward functions based on physical conditions. Our approach allows the model to dynamically adjust and improve the physical properties of generated videos, ensuring adherence to physical laws. Experimental results demonstrate that PhysAR can generate videos that are physically consistent.
title Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.15932