EgoForge: Goal-Directed Egocentric World Simulator

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shen, Yifan, Liu, Jiateng, Li, Xinzhuo, Liu, Yuanzhe, Li, Bingxuan, Yang, Houze, Jia, Wenqi, Li, Yijiang, Yu, Tianjiao, Rehg, James Matthew, Cao, Xu, Lourentzou, Ismini
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911532208619520
author Shen, Yifan
Liu, Jiateng
Li, Xinzhuo
Liu, Yuanzhe
Li, Bingxuan
Yang, Houze
Jia, Wenqi
Li, Yijiang
Yu, Tianjiao
Rehg, James Matthew
Cao, Xu
Lourentzou, Ismini
author_facet Shen, Yifan
Liu, Jiateng
Li, Xinzhuo
Liu, Yuanzhe
Li, Bingxuan
Yang, Houze
Jia, Wenqi
Li, Yijiang
Yu, Tianjiao
Rehg, James Matthew
Cao, Xu
Lourentzou, Ismini
contents Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends on latent human intent. Existing approaches either focus on hand-centric instructional synthesis with limited scene evolution, perform static view translation without modeling action dynamics, or rely on dense supervision, such as camera trajectories, long video prefixes, synchronized multicamera capture, etc. In this work, we introduce EgoForge, an egocentric goal-directed world simulator that generates coherent, first-person video rollouts from minimal static inputs: a single egocentric image, a high-level instruction, and an optional auxiliary exocentric view. To improve intent alignment and temporal consistency, we propose VideoDiffusionNFT, a trajectory-level reward-guided refinement that optimizes goal completion, temporal causality, scene consistency, and perceptual fidelity during diffusion sampling. Extensive experiments show EgoForge achieves consistent gains in semantic alignment, geometric stability, and motion fidelity over strong baselines, and robust performance in real-world smart-glasses experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2603_20169
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EgoForge: Goal-Directed Egocentric World Simulator
Shen, Yifan
Liu, Jiateng
Li, Xinzhuo
Liu, Yuanzhe
Li, Bingxuan
Yang, Houze
Jia, Wenqi
Li, Yijiang
Yu, Tianjiao
Rehg, James Matthew
Cao, Xu
Lourentzou, Ismini
Computer Vision and Pattern Recognition
Multimedia
Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends on latent human intent. Existing approaches either focus on hand-centric instructional synthesis with limited scene evolution, perform static view translation without modeling action dynamics, or rely on dense supervision, such as camera trajectories, long video prefixes, synchronized multicamera capture, etc. In this work, we introduce EgoForge, an egocentric goal-directed world simulator that generates coherent, first-person video rollouts from minimal static inputs: a single egocentric image, a high-level instruction, and an optional auxiliary exocentric view. To improve intent alignment and temporal consistency, we propose VideoDiffusionNFT, a trajectory-level reward-guided refinement that optimizes goal completion, temporal causality, scene consistency, and perceptual fidelity during diffusion sampling. Extensive experiments show EgoForge achieves consistent gains in semantic alignment, geometric stability, and motion fidelity over strong baselines, and robust performance in real-world smart-glasses experiments.
title EgoForge: Goal-Directed Egocentric World Simulator
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2603.20169