Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Weitai, Qu, Mengxue, Kini, Jyoti, Wei, Yunchao, Shah, Mubarak, Yan, Yan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917926537265152
author Kang, Weitai
Qu, Mengxue
Kini, Jyoti
Wei, Yunchao
Shah, Mubarak
Yan, Yan
author_facet Kang, Weitai
Qu, Mengxue
Kini, Jyoti
Wei, Yunchao
Shah, Mubarak
Yan, Yan
contents In real-life scenarios, humans seek out objects in the 3D world to fulfill their daily needs or intentions. This inspires us to introduce 3D intention grounding, a new task in 3D object detection employing RGB-D, based on human intention, such as "I want something to support my back". Closely related, 3D visual grounding focuses on understanding human reference. To achieve detection based on human intention, it relies on humans to observe the scene, reason out the target that aligns with their intention ("pillow" in this case), and finally provide a reference to the AI system, such as "A pillow on the couch". Instead, 3D intention grounding challenges AI agents to automatically observe, reason and detect the desired target solely based on human intention. To tackle this challenge, we introduce the new Intent3D dataset, consisting of 44,990 intention texts associated with 209 fine-grained classes from 1,042 scenes of the ScanNet dataset. We also establish several baselines based on different language-based 3D object detection models on our benchmark. Finally, we propose IntentNet, our unique approach, designed to tackle this intention-based detection problem. It focuses on three key aspects: intention understanding, reasoning to identify object candidates, and cascaded adaptive learning that leverages the intrinsic priority logic of different losses for multiple objective optimization. Project Page: https://weitaikang.github.io/Intent3D-webpage/
format Preprint
id arxiv_https___arxiv_org_abs_2405_18295
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention
Kang, Weitai
Qu, Mengxue
Kini, Jyoti
Wei, Yunchao
Shah, Mubarak
Yan, Yan
Computer Vision and Pattern Recognition
In real-life scenarios, humans seek out objects in the 3D world to fulfill their daily needs or intentions. This inspires us to introduce 3D intention grounding, a new task in 3D object detection employing RGB-D, based on human intention, such as "I want something to support my back". Closely related, 3D visual grounding focuses on understanding human reference. To achieve detection based on human intention, it relies on humans to observe the scene, reason out the target that aligns with their intention ("pillow" in this case), and finally provide a reference to the AI system, such as "A pillow on the couch". Instead, 3D intention grounding challenges AI agents to automatically observe, reason and detect the desired target solely based on human intention. To tackle this challenge, we introduce the new Intent3D dataset, consisting of 44,990 intention texts associated with 209 fine-grained classes from 1,042 scenes of the ScanNet dataset. We also establish several baselines based on different language-based 3D object detection models on our benchmark. Finally, we propose IntentNet, our unique approach, designed to tackle this intention-based detection problem. It focuses on three key aspects: intention understanding, reasoning to identify object candidates, and cascaded adaptive learning that leverages the intrinsic priority logic of different losses for multiple objective optimization. Project Page: https://weitaikang.github.io/Intent3D-webpage/
title Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.18295