GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ackermann, Jan, Cai, Shengqu, Deng, Boyang, Kuang, Zhengfei, Peng, Songyou, Wetzstein, Gordon
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917508093575168
author Ackermann, Jan
Cai, Shengqu
Deng, Boyang
Kuang, Zhengfei
Peng, Songyou
Wetzstein, Gordon
author_facet Ackermann, Jan
Cai, Shengqu
Deng, Boyang
Kuang, Zhengfei
Peng, Songyou
Wetzstein, Gordon
contents Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object deformation, texture drift, and non-rigid backgrounds under camera motion. Existing solutions either improve consistency as a byproduct, apply only to static scenes or realign the latent space of the model completely. We introduce a geometry-consistency reward that directly measures whether motion in a generated video is compatible with a coherent scene. Our key insight is that in physically consistent videos, background motion should be explainable by rigid camera-induced flow, while independently moving objects should preserve appearance identity along motion trajectories. We operationalize this using optical flow, depth--pose predictions, and feature-based correspondence to separate rigid and dynamic regions and evaluate their respective consistency. Integrating this reward with reinforcement fine-tuning transforms geometric consistency from an emergent property into an explicit optimization objective for video generators. The approach is model agnostic and applies to diverse dynamic scenes containing both camera and object motion. Experiments show substantial reductions in temporal geometric artifacts over strong baselines while preserving perceptual quality. Code and model weights are published.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18365
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation
Ackermann, Jan
Cai, Shengqu
Deng, Boyang
Kuang, Zhengfei
Peng, Songyou
Wetzstein, Gordon
Computer Vision and Pattern Recognition
Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object deformation, texture drift, and non-rigid backgrounds under camera motion. Existing solutions either improve consistency as a byproduct, apply only to static scenes or realign the latent space of the model completely. We introduce a geometry-consistency reward that directly measures whether motion in a generated video is compatible with a coherent scene. Our key insight is that in physically consistent videos, background motion should be explainable by rigid camera-induced flow, while independently moving objects should preserve appearance identity along motion trajectories. We operationalize this using optical flow, depth--pose predictions, and feature-based correspondence to separate rigid and dynamic regions and evaluate their respective consistency. Integrating this reward with reinforcement fine-tuning transforms geometric consistency from an emergent property into an explicit optimization objective for video generators. The approach is model agnostic and applies to diverse dynamic scenes containing both camera and object motion. Experiments show substantial reductions in temporal geometric artifacts over strong baselines while preserving perceptual quality. Code and model weights are published.
title GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.18365