Learning Robust 3D Representation from CLIP via Dual Denoising

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Luo, Shuqing, Qu, Bowen, Gao, Wei
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909235334348800
author Luo, Shuqing
Qu, Bowen
Gao, Wei
author_facet Luo, Shuqing
Qu, Bowen
Gao, Wei
contents In this paper, we explore a critical yet under-investigated issue: how to learn robust and well-generalized 3D representation from pre-trained vision language models such as CLIP. Previous works have demonstrated that cross-modal distillation can provide rich and useful knowledge for 3D data. However, like most deep learning models, the resultant 3D learning network is still vulnerable to adversarial attacks especially the iterative attack. In this work, we propose Dual Denoising, a novel framework for learning robust and well-generalized 3D representations from CLIP. It combines a denoising-based proxy task with a novel feature denoising network for 3D pre-training. Additionally, we propose utilizing parallel noise inference to enhance the generalization of point cloud features under cross domain settings. Experiments show that our model can effectively improve the representation learning performance and adversarial robustness of the 3D learning network under zero-shot settings without adversarial training. Our code is available at https://github.com/luoshuqing2001/Dual_Denoising.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00905
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Robust 3D Representation from CLIP via Dual Denoising
Luo, Shuqing
Qu, Bowen
Gao, Wei
Computer Vision and Pattern Recognition
In this paper, we explore a critical yet under-investigated issue: how to learn robust and well-generalized 3D representation from pre-trained vision language models such as CLIP. Previous works have demonstrated that cross-modal distillation can provide rich and useful knowledge for 3D data. However, like most deep learning models, the resultant 3D learning network is still vulnerable to adversarial attacks especially the iterative attack. In this work, we propose Dual Denoising, a novel framework for learning robust and well-generalized 3D representations from CLIP. It combines a denoising-based proxy task with a novel feature denoising network for 3D pre-training. Additionally, we propose utilizing parallel noise inference to enhance the generalization of point cloud features under cross domain settings. Experiments show that our model can effectively improve the representation learning performance and adversarial robustness of the 3D learning network under zero-shot settings without adversarial training. Our code is available at https://github.com/luoshuqing2001/Dual_Denoising.
title Learning Robust 3D Representation from CLIP via Dual Denoising
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.00905