Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Kaiqiang, Li, Gang, Zhou, Mingle, Li, Min, Han, Delong, Wan, Jin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915881539338240
author Li, Kaiqiang
Li, Gang
Zhou, Mingle
Li, Min
Han, Delong
Wan, Jin
author_facet Li, Kaiqiang
Li, Gang
Zhou, Mingle
Li, Min
Han, Delong
Wan, Jin
contents Zero-shot (ZS) 3D anomaly detection is crucial for reliable industrial inspection, as it enables detecting and localizing defects without requiring any target-category training data. Existing approaches render 3D point clouds into 2D images and leverage pre-trained Vision-Language Models (VLMs) for anomaly detection. However, such strategies inevitably discard geometric details and exhibit limited sensitivity to local anomalies. In this paper, we revisit intrinsic 3D representations and explore the potential of pre-trained Point-Language Models (PLMs) for ZS 3D anomaly detection. We propose BTP (Back To Point), a novel framework that effectively aligns 3D point cloud and textual embeddings. Specifically, BTP aligns multi-granularity patch features with textual representations for localized anomaly detection, while incorporating geometric descriptors to enhance sensitivity to structural anomalies. Furthermore, we introduce a joint representation learning strategy that leverages auxiliary point cloud data to improve robustness and enrich anomaly semantics. Extensive experiments on Real3D-AD and Anomaly-ShapeNet demonstrate that BTP achieves superior performance in ZS 3D anomaly detection. Code will be available at \href{https://github.com/wistful-8029/BTP-3DAD}{https://github.com/wistful-8029/BTP-3DAD}.
format Preprint
id arxiv_https___arxiv_org_abs_2603_21511
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection
Li, Kaiqiang
Li, Gang
Zhou, Mingle
Li, Min
Han, Delong
Wan, Jin
Computer Vision and Pattern Recognition
Zero-shot (ZS) 3D anomaly detection is crucial for reliable industrial inspection, as it enables detecting and localizing defects without requiring any target-category training data. Existing approaches render 3D point clouds into 2D images and leverage pre-trained Vision-Language Models (VLMs) for anomaly detection. However, such strategies inevitably discard geometric details and exhibit limited sensitivity to local anomalies. In this paper, we revisit intrinsic 3D representations and explore the potential of pre-trained Point-Language Models (PLMs) for ZS 3D anomaly detection. We propose BTP (Back To Point), a novel framework that effectively aligns 3D point cloud and textual embeddings. Specifically, BTP aligns multi-granularity patch features with textual representations for localized anomaly detection, while incorporating geometric descriptors to enhance sensitivity to structural anomalies. Furthermore, we introduce a joint representation learning strategy that leverages auxiliary point cloud data to improve robustness and enrich anomaly semantics. Extensive experiments on Real3D-AD and Anomaly-ShapeNet demonstrate that BTP achieves superior performance in ZS 3D anomaly detection. Code will be available at \href{https://github.com/wistful-8029/BTP-3DAD}{https://github.com/wistful-8029/BTP-3DAD}.
title Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.21511