TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Bu, Zheng, Yupeng, Li, Pengfei, Li, Weize, Zheng, Yuhang, Hu, Sujie, Liu, Xinyu, Zhu, Jinwei, Yan, Zhijie, Sun, Haiyang, Zhan, Kun, Jia, Peng, Long, Xiaoxiao, Chen, Yilun, Zhao, Hao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909218025504768
author Jin, Bu
Zheng, Yupeng
Li, Pengfei
Li, Weize
Zheng, Yuhang
Hu, Sujie
Liu, Xinyu
Zhu, Jinwei
Yan, Zhijie
Sun, Haiyang
Zhan, Kun
Jia, Peng
Long, Xiaoxiao
Chen, Yilun
Zhao, Hao
author_facet Jin, Bu
Zheng, Yupeng
Li, Pengfei
Li, Weize
Zheng, Yuhang
Hu, Sujie
Liu, Xinyu
Zhu, Jinwei
Yan, Zhijie
Sun, Haiyang
Zhan, Kun
Jia, Peng
Long, Xiaoxiao
Chen, Yilun
Zhao, Hao
contents 3D dense captioning stands as a cornerstone in achieving a comprehensive understanding of 3D scenes through natural language. It has recently witnessed remarkable achievements, particularly in indoor settings. However, the exploration of 3D dense captioning in outdoor scenes is hindered by two major challenges: 1) the domain gap between indoor and outdoor scenes, such as dynamics and sparse visual inputs, makes it difficult to directly adapt existing indoor methods; 2) the lack of data with comprehensive box-caption pair annotations specifically tailored for outdoor scenes. To this end, we introduce the new task of outdoor 3D dense captioning. As input, we assume a LiDAR point cloud and a set of RGB images captured by the panoramic camera rig. The expected output is a set of object boxes with captions. To tackle this task, we propose the TOD3Cap network, which leverages the BEV representation to generate object box proposals and integrates Relation Q-Former with LLaMA-Adapter to generate rich captions for these objects. We also introduce the TOD3Cap dataset, the largest one to our knowledge for 3D dense captioning in outdoor scenes, which contains 2.3M descriptions of 64.3K outdoor objects from 850 scenes. Notably, our TOD3Cap network can effectively localize and caption 3D objects in outdoor scenes, which outperforms baseline methods by a significant margin (+9.6 CiDEr@0.5IoU). Code, data, and models are publicly available at https://github.com/jxbbb/TOD3Cap.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19589
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
Jin, Bu
Zheng, Yupeng
Li, Pengfei
Li, Weize
Zheng, Yuhang
Hu, Sujie
Liu, Xinyu
Zhu, Jinwei
Yan, Zhijie
Sun, Haiyang
Zhan, Kun
Jia, Peng
Long, Xiaoxiao
Chen, Yilun
Zhao, Hao
Computer Vision and Pattern Recognition
3D dense captioning stands as a cornerstone in achieving a comprehensive understanding of 3D scenes through natural language. It has recently witnessed remarkable achievements, particularly in indoor settings. However, the exploration of 3D dense captioning in outdoor scenes is hindered by two major challenges: 1) the domain gap between indoor and outdoor scenes, such as dynamics and sparse visual inputs, makes it difficult to directly adapt existing indoor methods; 2) the lack of data with comprehensive box-caption pair annotations specifically tailored for outdoor scenes. To this end, we introduce the new task of outdoor 3D dense captioning. As input, we assume a LiDAR point cloud and a set of RGB images captured by the panoramic camera rig. The expected output is a set of object boxes with captions. To tackle this task, we propose the TOD3Cap network, which leverages the BEV representation to generate object box proposals and integrates Relation Q-Former with LLaMA-Adapter to generate rich captions for these objects. We also introduce the TOD3Cap dataset, the largest one to our knowledge for 3D dense captioning in outdoor scenes, which contains 2.3M descriptions of 64.3K outdoor objects from 850 scenes. Notably, our TOD3Cap network can effectively localize and caption 3D objects in outdoor scenes, which outperforms baseline methods by a significant margin (+9.6 CiDEr@0.5IoU). Code, data, and models are publicly available at https://github.com/jxbbb/TOD3Cap.
title TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.19589