Multi-Task Domain Adaptation for Language Grounding with 3D Objects

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Penglei, Song, Yaoxian, Pan, Xinglin, Dong, Peijie, Yang, Xiaofei, Wang, Qiang, Li, Zhixu, Li, Tiefeng, Chu, Xiaowen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913417401466880
author Sun, Penglei
Song, Yaoxian
Pan, Xinglin
Dong, Peijie
Yang, Xiaofei
Wang, Qiang
Li, Zhixu
Li, Tiefeng
Chu, Xiaowen
author_facet Sun, Penglei
Song, Yaoxian
Pan, Xinglin
Dong, Peijie
Yang, Xiaofei
Wang, Qiang
Li, Zhixu
Li, Tiefeng
Chu, Xiaowen
contents The existing works on object-level language grounding with 3D objects mostly focus on improving performance by utilizing the off-the-shelf pre-trained models to capture features, such as viewpoint selection or geometric priors. However, they have failed to consider exploring the cross-modal representation of language-vision alignment in the cross-domain field. To answer this problem, we propose a novel method called Domain Adaptation for Language Grounding (DA4LG) with 3D objects. Specifically, the proposed DA4LG consists of a visual adapter module with multi-task learning to realize vision-language alignment by comprehensive multimodal feature representation. Experimental results demonstrate that DA4LG competitively performs across visual and non-visual language descriptions, independent of the completeness of observation. DA4LG achieves state-of-the-art performance in the single-view setting and multi-view setting with the accuracy of 83.8% and 86.8% respectively in the language grounding benchmark SNARE. The simulation experiments show the well-practical and generalized performance of DA4LG compared to the existing methods. Our project is available at https://sites.google.com/view/da4lg.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02846
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Task Domain Adaptation for Language Grounding with 3D Objects
Sun, Penglei
Song, Yaoxian
Pan, Xinglin
Dong, Peijie
Yang, Xiaofei
Wang, Qiang
Li, Zhixu
Li, Tiefeng
Chu, Xiaowen
Computer Vision and Pattern Recognition
The existing works on object-level language grounding with 3D objects mostly focus on improving performance by utilizing the off-the-shelf pre-trained models to capture features, such as viewpoint selection or geometric priors. However, they have failed to consider exploring the cross-modal representation of language-vision alignment in the cross-domain field. To answer this problem, we propose a novel method called Domain Adaptation for Language Grounding (DA4LG) with 3D objects. Specifically, the proposed DA4LG consists of a visual adapter module with multi-task learning to realize vision-language alignment by comprehensive multimodal feature representation. Experimental results demonstrate that DA4LG competitively performs across visual and non-visual language descriptions, independent of the completeness of observation. DA4LG achieves state-of-the-art performance in the single-view setting and multi-view setting with the accuracy of 83.8% and 86.8% respectively in the language grounding benchmark SNARE. The simulation experiments show the well-practical and generalized performance of DA4LG compared to the existing methods. Our project is available at https://sites.google.com/view/da4lg.
title Multi-Task Domain Adaptation for Language Grounding with 3D Objects
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.02846