Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhichuan, Zhou, Yang, Liu, Zhe, Yu, Rui, Bai, Song, Wang, Yulong, He, Xinwei, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
by: He, Xinwei, et al.
Published: (2026)
by: He, Xinwei, et al.
Published: (2026)
TeDA: Boosting Vision-Lanuage Models for Zero-Shot 3D Object Retrieval via Testing-time Distribution Alignment
by: Wang, Zhichuan, et al.
Published: (2025)
by: Wang, Zhichuan, et al.
Published: (2025)
Anomaly Detection by Adapting a pre-trained Vision Language Model
by: Cai, Yuxuan, et al.
Published: (2024)
by: Cai, Yuxuan, et al.
Published: (2024)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2025)
by: Gao, Bin-Bin, et al.
Published: (2025)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
by: Song, Dan, et al.
Published: (2023)
by: Song, Dan, et al.
Published: (2023)
M3: A Multi-Task Mixed-Objective Learning Framework for Open-Domain Multi-Hop Dense Sentence Retrieval
by: Bai, Yang, et al.
Published: (2024)
by: Bai, Yang, et al.
Published: (2024)
DialCLIP: Empowering CLIP as Multi-Modal Dialog Retriever
by: Yin, Zhichao, et al.
Published: (2024)
by: Yin, Zhichao, et al.
Published: (2024)
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
by: Zhu, Wenqi, et al.
Published: (2024)
by: Zhu, Wenqi, et al.
Published: (2024)
LION: Linear Group RNN for 3D Object Detection in Point Clouds
by: Liu, Zhe, et al.
Published: (2024)
by: Liu, Zhe, et al.
Published: (2024)
OPEN: Object-wise Position Embedding for Multi-view 3D Object Detection
by: Hou, Jinghua, et al.
Published: (2024)
by: Hou, Jinghua, et al.
Published: (2024)
SAM3D: Zero-Shot 3D Object Detection via Segment Anything Model
by: Zhang, Dingyuan, et al.
Published: (2023)
by: Zhang, Dingyuan, et al.
Published: (2023)
A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
Toward Open Vocabulary Aerial Object Detection with CLIP-Activated Student-Teacher Learning
by: Li, Yan, et al.
Published: (2023)
by: Li, Yan, et al.
Published: (2023)
CS3D: An Efficient Facial Expression Recognition via Event Vision
by: Wang, Zhe, et al.
Published: (2025)
by: Wang, Zhe, et al.
Published: (2025)
Align then Adapt: Rethinking Parameter-Efficient Transfer Learning in 4D Perception
by: Sun, Yiding, et al.
Published: (2026)
by: Sun, Yiding, et al.
Published: (2026)
Towards 3D VR-Sketch to 3D Shape Retrieval
by: Luo, Ling, et al.
Published: (2022)
by: Luo, Ling, et al.
Published: (2022)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2023)
by: Xiao, Linhui, et al.
Published: (2023)
Boosting Instance Awareness via Cross-View Correlation with 4D Radar and Camera for 3D Object Detection
by: Bai, Xiaokai, et al.
Published: (2026)
by: Bai, Xiaokai, et al.
Published: (2026)
RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question Answering
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
by: Song, Yehun, et al.
Published: (2025)
by: Song, Yehun, et al.
Published: (2025)
CUS3D :CLIP-based Unsupervised 3D Segmentation via Object-level Denoise
by: Yu, Fuyang, et al.
Published: (2024)
by: Yu, Fuyang, et al.
Published: (2024)
You Only Look Bottom-Up for Monocular 3D Object Detection
by: Xiong, Kaixin, et al.
Published: (2024)
by: Xiong, Kaixin, et al.
Published: (2024)
Towards Camera Open-set 3D Object Detection for Autonomous Driving Scenarios
by: He, Zhuolin, et al.
Published: (2024)
by: He, Zhuolin, et al.
Published: (2024)
FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Incremental Object Detection with CLIP
by: Huang, Ziyue, et al.
Published: (2023)
by: Huang, Ziyue, et al.
Published: (2023)
Thermal Resistance of 1D Micron‐Wide Interface Under External Force Modulation
by: Amin Karamati, et al.
Published: (2025)
by: Amin Karamati, et al.
Published: (2025)
SEED: A Simple and Effective 3D DETR in Point Clouds
by: Liu, Zhe, et al.
Published: (2024)
by: Liu, Zhe, et al.
Published: (2024)
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation
by: Bai, Sule, et al.
Published: (2024)
by: Bai, Sule, et al.
Published: (2024)
PartGLEE: A Foundation Model for Recognizing and Parsing Any Objects
by: Li, Junyi, et al.
Published: (2024)
by: Li, Junyi, et al.
Published: (2024)
Keypoint-based Dynamic Object 6-DoF Pose Tracking via Event Camera
by: Wang, Zhe, et al.
Published: (2026)
by: Wang, Zhe, et al.
Published: (2026)
Open-vocabulary vs. Closed-set: Best Practice for Few-shot Object Detection Considering Text Describability
by: Hosoya, Yusuke, et al.
Published: (2024)
by: Hosoya, Yusuke, et al.
Published: (2024)
Adapting Segment Anything Model for Unseen Object Instance Segmentation
by: Cao, Rui, et al.
Published: (2024)
by: Cao, Rui, et al.
Published: (2024)
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task
by: Asgarov, Ali, et al.
Published: (2024)
by: Asgarov, Ali, et al.
Published: (2024)
Exploring Open-Vocabulary Object Recognition in Images using CLIP
by: Chen, Wei Yu, et al.
Published: (2026)
by: Chen, Wei Yu, et al.
Published: (2026)
URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Approaching Code Search for Python as a Translation Retrieval Problem with Dual Encoders
by: Khan, Monoshiz Mahbub, et al.
Published: (2024)
by: Khan, Monoshiz Mahbub, et al.
Published: (2024)
Omni-AD: Learning to Reconstruct Global and Local Features for Multi-class Anomaly Detection
by: Quan, Jiajie, et al.
Published: (2025)
by: Quan, Jiajie, et al.
Published: (2025)
Adapting Segment Anything Model for 3D Brain Tumor Segmentation With Missing Modalities
by: Xiaoliang Lei, et al.
Published: (2024)
by: Xiaoliang Lei, et al.
Published: (2024)
Implementing Multimodal Hardware Security with 2D α‐In 2 Se 3 Ferroelectric Transistor
by: Xinwei Zhang, et al.
Published: (2025)
by: Xinwei Zhang, et al.
Published: (2025)
Role of $4S$-$3D$ mixing in explaining the $ω$-like $Y(2119)$ observed in $e^+e^-\toρπ$ and $ρ(1450)π$
by: Bai, Zi-Yue, et al.
Published: (2025)
by: Bai, Zi-Yue, et al.
Published: (2025)
Similar Items
-
DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
by: He, Xinwei, et al.
Published: (2026) -
TeDA: Boosting Vision-Lanuage Models for Zero-Shot 3D Object Retrieval via Testing-time Distribution Alignment
by: Wang, Zhichuan, et al.
Published: (2025) -
Anomaly Detection by Adapting a pre-trained Vision Language Model
by: Cai, Yuxuan, et al.
Published: (2024) -
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2025) -
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
by: Song, Dan, et al.
Published: (2023)