WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ying, Qijun, Hu, Zhongyuan, Zhang, Rui, Li, Ronghui, Lu, Yu, Zeng, Zijiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908519946518528
author Ying, Qijun
Hu, Zhongyuan
Zhang, Rui
Li, Ronghui
Lu, Yu
Zeng, Zijiao
author_facet Ying, Qijun
Hu, Zhongyuan
Zhang, Rui
Li, Ronghui
Lu, Yu
Zeng, Zijiao
contents Global human motion reconstruction from in-the-wild monocular videos is increasingly demanded across VR, graphics, and robotics applications, yet requires accurate mapping of human poses from camera to world coordinates-a task challenged by depth ambiguity, motion ambiguity, and the entanglement between camera and human movements. While human-motion-centric approaches excel in preserving motion details and physical plausibility, they suffer from two critical limitations: insufficient exploitation of camera orientation information and ineffective integration of camera translation cues. We present WATCH (World-aware Allied Trajectory and pose reconstruction for Camera and Human), a unified framework addressing both challenges. Our approach introduces an analytical heading angle decomposition technique that offers superior efficiency and extensibility compared to existing geometric methods. Additionally, we design a camera trajectory integration mechanism inspired by world models, providing an effective pathway for leveraging camera translation information beyond naive hard-decoding approaches. Through experiments on in-the-wild benchmarks, WATCH achieves state-of-the-art performance in end-to-end trajectory reconstruction. Our work demonstrates the effectiveness of jointly modeling camera-human motion relationships and offers new insights for addressing the long-standing challenge of camera translation integration in global human motion reconstruction. The code will be available publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04600
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human
Ying, Qijun
Hu, Zhongyuan
Zhang, Rui
Li, Ronghui
Lu, Yu
Zeng, Zijiao
Computer Vision and Pattern Recognition
Global human motion reconstruction from in-the-wild monocular videos is increasingly demanded across VR, graphics, and robotics applications, yet requires accurate mapping of human poses from camera to world coordinates-a task challenged by depth ambiguity, motion ambiguity, and the entanglement between camera and human movements. While human-motion-centric approaches excel in preserving motion details and physical plausibility, they suffer from two critical limitations: insufficient exploitation of camera orientation information and ineffective integration of camera translation cues. We present WATCH (World-aware Allied Trajectory and pose reconstruction for Camera and Human), a unified framework addressing both challenges. Our approach introduces an analytical heading angle decomposition technique that offers superior efficiency and extensibility compared to existing geometric methods. Additionally, we design a camera trajectory integration mechanism inspired by world models, providing an effective pathway for leveraging camera translation information beyond naive hard-decoding approaches. Through experiments on in-the-wild benchmarks, WATCH achieves state-of-the-art performance in end-to-end trajectory reconstruction. Our work demonstrates the effectiveness of jointly modeling camera-human motion relationships and offers new insights for addressing the long-standing challenge of camera translation integration in global human motion reconstruction. The code will be available publicly.
title WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.04600