Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guan, Runwei, Zhang, Ruixiao, Ouyang, Ningwei, Liu, Jianan, Man, Ka Lok, Cai, Xiaohao, Xu, Ming, Smith, Jeremy, Lim, Eng Gee, Yue, Yutao, Xiong, Hui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910819510386688
author Guan, Runwei
Zhang, Ruixiao
Ouyang, Ningwei
Liu, Jianan
Man, Ka Lok
Cai, Xiaohao
Xu, Ming
Smith, Jeremy
Lim, Eng Gee
Yue, Yutao
Xiong, Hui
author_facet Guan, Runwei
Zhang, Ruixiao
Ouyang, Ningwei
Liu, Jianan
Man, Ka Lok
Cai, Xiaohao
Xu, Ming
Smith, Jeremy
Lim, Eng Gee
Yue, Yutao
Xiong, Hui
contents Embodied perception is essential for intelligent vehicles and robots in interactive environmental understanding. However, these advancements primarily focus on vision, with limited attention given to using 3D modeling sensors, restricting a comprehensive understanding of objects in response to prompts containing qualitative and quantitative queries. Recently, as a promising automotive sensor with affordable cost, 4D millimeter-wave radars provide denser point clouds than conventional radars and perceive both semantic and physical characteristics of objects, thereby enhancing the reliability of perception systems. To foster the development of natural language-driven context understanding in radar scenes for 3D visual grounding, we construct the first dataset, Talk2Radar, which bridges these two modalities for 3D Referring Expression Comprehension (REC). Talk2Radar contains 8,682 referring prompt samples with 20,558 referred objects. Moreover, we propose a novel model, T-RadarNet, for 3D REC on point clouds, achieving State-Of-The-Art (SOTA) performance on the Talk2Radar dataset compared to counterparts. Deformable-FPN and Gated Graph Fusion are meticulously designed for efficient point cloud feature modeling and cross-modal fusion between radar and text features, respectively. Comprehensive experiments provide deep insights into radar-based 3D REC. We release our project at https://github.com/GuanRunwei/Talk2Radar.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12821
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension
Guan, Runwei
Zhang, Ruixiao
Ouyang, Ningwei
Liu, Jianan
Man, Ka Lok
Cai, Xiaohao
Xu, Ming
Smith, Jeremy
Lim, Eng Gee
Yue, Yutao
Xiong, Hui
Robotics
Computer Vision and Pattern Recognition
Embodied perception is essential for intelligent vehicles and robots in interactive environmental understanding. However, these advancements primarily focus on vision, with limited attention given to using 3D modeling sensors, restricting a comprehensive understanding of objects in response to prompts containing qualitative and quantitative queries. Recently, as a promising automotive sensor with affordable cost, 4D millimeter-wave radars provide denser point clouds than conventional radars and perceive both semantic and physical characteristics of objects, thereby enhancing the reliability of perception systems. To foster the development of natural language-driven context understanding in radar scenes for 3D visual grounding, we construct the first dataset, Talk2Radar, which bridges these two modalities for 3D Referring Expression Comprehension (REC). Talk2Radar contains 8,682 referring prompt samples with 20,558 referred objects. Moreover, we propose a novel model, T-RadarNet, for 3D REC on point clouds, achieving State-Of-The-Art (SOTA) performance on the Talk2Radar dataset compared to counterparts. Deformable-FPN and Gated Graph Fusion are meticulously designed for efficient point cloud feature modeling and cross-modal fusion between radar and text features, respectively. Comprehensive experiments provide deep insights into radar-based 3D REC. We release our project at https://github.com/GuanRunwei/Talk2Radar.
title Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.12821