Dynamic Camera Poses and Where to Find Them
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908336815865856 |
|---|---|
| author | Rockwell, Chris Tung, Joseph Lin, Tsung-Yi Liu, Ming-Yu Fouhey, David F. Lin, Chen-Hsuan |
| author_facet | Rockwell, Chris Tung, Joseph Lin, Tsung-Yi Liu, Ming-Yu Fouhey, David F. Lin, Chen-Hsuan |
| contents | Annotating camera poses on dynamic Internet videos at scale is critical for advancing fields like realistic video generation and simulation. However, collecting such a dataset is difficult, as most Internet videos are unsuitable for pose estimation. Furthermore, annotating dynamic Internet videos present significant challenges even for state-of-theart methods. In this paper, we introduce DynPose-100K, a large-scale dataset of dynamic Internet videos annotated with camera poses. Our collection pipeline addresses filtering using a carefully combined set of task-specific and generalist models. For pose estimation, we combine the latest techniques of point tracking, dynamic masking, and structure-from-motion to achieve improvements over the state-of-the-art approaches. Our analysis and experiments demonstrate that DynPose-100K is both large-scale and diverse across several key attributes, opening up avenues for advancements in various downstream applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_17788 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Dynamic Camera Poses and Where to Find Them Rockwell, Chris Tung, Joseph Lin, Tsung-Yi Liu, Ming-Yu Fouhey, David F. Lin, Chen-Hsuan Computer Vision and Pattern Recognition Annotating camera poses on dynamic Internet videos at scale is critical for advancing fields like realistic video generation and simulation. However, collecting such a dataset is difficult, as most Internet videos are unsuitable for pose estimation. Furthermore, annotating dynamic Internet videos present significant challenges even for state-of-theart methods. In this paper, we introduce DynPose-100K, a large-scale dataset of dynamic Internet videos annotated with camera poses. Our collection pipeline addresses filtering using a carefully combined set of task-specific and generalist models. For pose estimation, we combine the latest techniques of point tracking, dynamic masking, and structure-from-motion to achieve improvements over the state-of-the-art approaches. Our analysis and experiments demonstrate that DynPose-100K is both large-scale and diverse across several key attributes, opening up avenues for advancements in various downstream applications. |
| title | Dynamic Camera Poses and Where to Find Them |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2504.17788 |