Dynamic Camera Poses and Where to Find Them

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rockwell, Chris, Tung, Joseph, Lin, Tsung-Yi, Liu, Ming-Yu, Fouhey, David F., Lin, Chen-Hsuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908336815865856
author Rockwell, Chris
Tung, Joseph
Lin, Tsung-Yi
Liu, Ming-Yu
Fouhey, David F.
Lin, Chen-Hsuan
author_facet Rockwell, Chris
Tung, Joseph
Lin, Tsung-Yi
Liu, Ming-Yu
Fouhey, David F.
Lin, Chen-Hsuan
contents Annotating camera poses on dynamic Internet videos at scale is critical for advancing fields like realistic video generation and simulation. However, collecting such a dataset is difficult, as most Internet videos are unsuitable for pose estimation. Furthermore, annotating dynamic Internet videos present significant challenges even for state-of-theart methods. In this paper, we introduce DynPose-100K, a large-scale dataset of dynamic Internet videos annotated with camera poses. Our collection pipeline addresses filtering using a carefully combined set of task-specific and generalist models. For pose estimation, we combine the latest techniques of point tracking, dynamic masking, and structure-from-motion to achieve improvements over the state-of-the-art approaches. Our analysis and experiments demonstrate that DynPose-100K is both large-scale and diverse across several key attributes, opening up avenues for advancements in various downstream applications.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17788
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Camera Poses and Where to Find Them
Rockwell, Chris
Tung, Joseph
Lin, Tsung-Yi
Liu, Ming-Yu
Fouhey, David F.
Lin, Chen-Hsuan
Computer Vision and Pattern Recognition
Annotating camera poses on dynamic Internet videos at scale is critical for advancing fields like realistic video generation and simulation. However, collecting such a dataset is difficult, as most Internet videos are unsuitable for pose estimation. Furthermore, annotating dynamic Internet videos present significant challenges even for state-of-theart methods. In this paper, we introduce DynPose-100K, a large-scale dataset of dynamic Internet videos annotated with camera poses. Our collection pipeline addresses filtering using a carefully combined set of task-specific and generalist models. For pose estimation, we combine the latest techniques of point tracking, dynamic masking, and structure-from-motion to achieve improvements over the state-of-the-art approaches. Our analysis and experiments demonstrate that DynPose-100K is both large-scale and diverse across several key attributes, opening up avenues for advancements in various downstream applications.
title Dynamic Camera Poses and Where to Find Them
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.17788