From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mahdi, Mohammad, Savov, Nedko, Paudel, Danda Pani, Van Gool, Luc
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915938607038464
author Mahdi, Mohammad
Savov, Nedko
Paudel, Danda Pani
Van Gool, Luc
author_facet Mahdi, Mohammad
Savov, Nedko
Paudel, Danda Pani
Van Gool, Luc
contents Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired supervision is available, synchronized exo-ego data inherently introduces substantial spatio-temporal and geometric discontinuities, violating the smooth-motion assumptions of standard video generation benchmarks. We identify this synchronization-induced jump as the central challenge and propose Syn2Seq-Forcing, a sequential formulation that interpolates between the source and target videos to form a single continuous signal. By reframing Exo2Ego as sequential signal modeling rather than a conventional condition-output task, our approach enables diffusion-based sequence models, e.g. Diffusion Forcing Transformers (DFoT), to capture coherent transitions across frames more effectively. Empirically, we show that interpolating only the videos, without performing pose interpolation already produces significant improvements, emphasizing that the dominant difficulty arises from spatio-temporal discontinuities. Beyond immediate performance gains, this formulation establishes a general and flexible framework capable of unifying both Exo2Ego and Ego2Exo generation within a single continuous sequence model, providing a principled foundation for future research in cross-view video synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13793
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
Mahdi, Mohammad
Savov, Nedko
Paudel, Danda Pani
Van Gool, Luc
Computer Vision and Pattern Recognition
Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired supervision is available, synchronized exo-ego data inherently introduces substantial spatio-temporal and geometric discontinuities, violating the smooth-motion assumptions of standard video generation benchmarks. We identify this synchronization-induced jump as the central challenge and propose Syn2Seq-Forcing, a sequential formulation that interpolates between the source and target videos to form a single continuous signal. By reframing Exo2Ego as sequential signal modeling rather than a conventional condition-output task, our approach enables diffusion-based sequence models, e.g. Diffusion Forcing Transformers (DFoT), to capture coherent transitions across frames more effectively. Empirically, we show that interpolating only the videos, without performing pose interpolation already produces significant improvements, emphasizing that the dominant difficulty arises from spatio-temporal discontinuities. Beyond immediate performance gains, this formulation establishes a general and flexible framework capable of unifying both Exo2Ego and Ego2Exo generation within a single continuous sequence model, providing a principled foundation for future research in cross-view video synthesis.
title From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.13793