Ego3DT: Tracking Every 3D Object in Ego-centric Videos
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866910645303115776 |
|---|---|
| author | Hao, Shengyu Chai, Wenhao Zhao, Zhonghan Sun, Meiqi Hu, Wendi Zhou, Jieyang Zhao, Yixian Li, Qi Wang, Yizhou Li, Xi Wang, Gaoang |
| author_facet | Hao, Shengyu Chai, Wenhao Zhao, Zhonghan Sun, Meiqi Hu, Wendi Zhou, Jieyang Zhao, Yixian Li, Qi Wang, Yizhou Li, Xi Wang, Gaoang |
| contents | The growing interest in embodied intelligence has brought ego-centric perspectives to contemporary research. One significant challenge within this realm is the accurate localization and tracking of objects in ego-centric videos, primarily due to the substantial variability in viewing angles. Addressing this issue, this paper introduces a novel zero-shot approach for the 3D reconstruction and tracking of all objects from the ego-centric video. We present Ego3DT, a novel framework that initially identifies and extracts detection and segmentation information of objects within the ego environment. Utilizing information from adjacent video frames, Ego3DT dynamically constructs a 3D scene of the ego view using a pre-trained 3D scene reconstruction model. Additionally, we have innovated a dynamic hierarchical association mechanism for creating stable 3D tracking trajectories of objects in ego-centric videos. Moreover, the efficacy of our approach is corroborated by extensive experiments on two newly compiled datasets, with 1.04x - 2.90x in HOTA, showcasing the robustness and accuracy of our method in diverse ego-centric scenarios. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_08530 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Ego3DT: Tracking Every 3D Object in Ego-centric Videos Hao, Shengyu Chai, Wenhao Zhao, Zhonghan Sun, Meiqi Hu, Wendi Zhou, Jieyang Zhao, Yixian Li, Qi Wang, Yizhou Li, Xi Wang, Gaoang Computer Vision and Pattern Recognition Multimedia The growing interest in embodied intelligence has brought ego-centric perspectives to contemporary research. One significant challenge within this realm is the accurate localization and tracking of objects in ego-centric videos, primarily due to the substantial variability in viewing angles. Addressing this issue, this paper introduces a novel zero-shot approach for the 3D reconstruction and tracking of all objects from the ego-centric video. We present Ego3DT, a novel framework that initially identifies and extracts detection and segmentation information of objects within the ego environment. Utilizing information from adjacent video frames, Ego3DT dynamically constructs a 3D scene of the ego view using a pre-trained 3D scene reconstruction model. Additionally, we have innovated a dynamic hierarchical association mechanism for creating stable 3D tracking trajectories of objects in ego-centric videos. Moreover, the efficacy of our approach is corroborated by extensive experiments on two newly compiled datasets, with 1.04x - 2.90x in HOTA, showcasing the robustness and accuracy of our method in diverse ego-centric scenarios. |
| title | Ego3DT: Tracking Every 3D Object in Ego-centric Videos |
| topic | Computer Vision and Pattern Recognition Multimedia |
| url | https://arxiv.org/abs/2410.08530 |