Multi-branch Collaborative Learning Network for 3D Visual Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Zhipeng, Ma, Yiwei, Lin, Zhekai, Ji, Jiayi, Zheng, Xiawu, Sun, Xiaoshuai, Ji, Rongrong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913424305291264
author Qian, Zhipeng
Ma, Yiwei
Lin, Zhekai
Ji, Jiayi
Zheng, Xiawu
Sun, Xiaoshuai
Ji, Rongrong
author_facet Qian, Zhipeng
Ma, Yiwei
Lin, Zhekai
Ji, Jiayi
Zheng, Xiawu
Sun, Xiaoshuai
Ji, Rongrong
contents 3D referring expression comprehension (3DREC) and segmentation (3DRES) have overlapping objectives, indicating their potential for collaboration. However, existing collaborative approaches predominantly depend on the results of one task to make predictions for the other, limiting effective collaboration. We argue that employing separate branches for 3DREC and 3DRES tasks enhances the model's capacity to learn specific information for each task, enabling them to acquire complementary knowledge. Thus, we propose the MCLN framework, which includes independent branches for 3DREC and 3DRES tasks. This enables dedicated exploration of each task and effective coordination between the branches. Furthermore, to facilitate mutual reinforcement between these branches, we introduce a Relative Superpoint Aggregation (RSA) module and an Adaptive Soft Alignment (ASA) module. These modules significantly contribute to the precise alignment of prediction results from the two branches, directing the module to allocate increased attention to key positions. Comprehensive experimental evaluation demonstrates that our proposed method achieves state-of-the-art performance on both the 3DREC and 3DRES tasks, with an increase of 2.05% in Acc@0.5 for 3DREC and 3.96% in mIoU for 3DRES.
format Preprint
id arxiv_https___arxiv_org_abs_2407_05363
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-branch Collaborative Learning Network for 3D Visual Grounding
Qian, Zhipeng
Ma, Yiwei
Lin, Zhekai
Ji, Jiayi
Zheng, Xiawu
Sun, Xiaoshuai
Ji, Rongrong
Computer Vision and Pattern Recognition
3D referring expression comprehension (3DREC) and segmentation (3DRES) have overlapping objectives, indicating their potential for collaboration. However, existing collaborative approaches predominantly depend on the results of one task to make predictions for the other, limiting effective collaboration. We argue that employing separate branches for 3DREC and 3DRES tasks enhances the model's capacity to learn specific information for each task, enabling them to acquire complementary knowledge. Thus, we propose the MCLN framework, which includes independent branches for 3DREC and 3DRES tasks. This enables dedicated exploration of each task and effective coordination between the branches. Furthermore, to facilitate mutual reinforcement between these branches, we introduce a Relative Superpoint Aggregation (RSA) module and an Adaptive Soft Alignment (ASA) module. These modules significantly contribute to the precise alignment of prediction results from the two branches, directing the module to allocate increased attention to key positions. Comprehensive experimental evaluation demonstrates that our proposed method achieves state-of-the-art performance on both the 3DREC and 3DRES tasks, with an increase of 2.05% in Acc@0.5 for 3DREC and 3.96% in mIoU for 3DRES.
title Multi-branch Collaborative Learning Network for 3D Visual Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.05363