Multimodal Robust Prompt Distillation for 3D Point Cloud Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Xiang, Lu, Liming, Zheng, Xu, Du, Anan, Zhou, Yongbin, Pang, Shuchao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917106653593600
author Gu, Xiang
Lu, Liming
Zheng, Xu
Du, Anan
Zhou, Yongbin
Pang, Shuchao
author_facet Gu, Xiang
Lu, Liming
Zheng, Xu
Du, Anan
Zhou, Yongbin
Pang, Shuchao
contents Adversarial attacks pose a significant threat to learning-based 3D point cloud models, critically undermining their reliability in security-sensitive applications. Existing defense methods often suffer from (1) high computational overhead and (2) poor generalization ability across diverse attack types. To bridge these gaps, we propose a novel yet efficient teacher-student framework, namely Multimodal Robust Prompt Distillation (MRPD) for distilling robust 3D point cloud model. It learns lightweight prompts by aligning student point cloud model's features with robust embeddings from three distinct teachers: a vision model processing depth projections, a high-performance 3D model, and a text encoder. To ensure a reliable knowledge transfer, this distillation is guided by a confidence-gated mechanism which dynamically balances the contribution of all input modalities. Notably, since the distillation is all during the training stage, there is no additional computational cost at inference. Extensive experiments demonstrate that MRPD substantially outperforms state-of-the-art defense methods against a wide range of white-box and black-box attacks, while even achieving better performance on clean data. Our work presents a new, practical paradigm for building robust 3D vision systems by efficiently harnessing multimodal knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21574
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Robust Prompt Distillation for 3D Point Cloud Models
Gu, Xiang
Lu, Liming
Zheng, Xu
Du, Anan
Zhou, Yongbin
Pang, Shuchao
Computer Vision and Pattern Recognition
Artificial Intelligence
Adversarial attacks pose a significant threat to learning-based 3D point cloud models, critically undermining their reliability in security-sensitive applications. Existing defense methods often suffer from (1) high computational overhead and (2) poor generalization ability across diverse attack types. To bridge these gaps, we propose a novel yet efficient teacher-student framework, namely Multimodal Robust Prompt Distillation (MRPD) for distilling robust 3D point cloud model. It learns lightweight prompts by aligning student point cloud model's features with robust embeddings from three distinct teachers: a vision model processing depth projections, a high-performance 3D model, and a text encoder. To ensure a reliable knowledge transfer, this distillation is guided by a confidence-gated mechanism which dynamically balances the contribution of all input modalities. Notably, since the distillation is all during the training stage, there is no additional computational cost at inference. Extensive experiments demonstrate that MRPD substantially outperforms state-of-the-art defense methods against a wide range of white-box and black-box attacks, while even achieving better performance on clean data. Our work presents a new, practical paradigm for building robust 3D vision systems by efficiently harnessing multimodal knowledge.
title Multimodal Robust Prompt Distillation for 3D Point Cloud Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.21574