From Points to Clouds: Learning Robust Semantic Distributions for Multi-modal Prompts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Weiran, Liu, Yeqiang, Wei, Yijie, Han, Mina, Liu, Xin, Li, Zhenbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911292040675328
author Li, Weiran
Liu, Yeqiang
Wei, Yijie
Han, Mina
Liu, Xin
Li, Zhenbo
author_facet Li, Weiran
Liu, Yeqiang
Wei, Yijie
Han, Mina
Liu, Xin
Li, Zhenbo
contents Multimodal Prompt Learning (MPL) has emerged as a pivotal technique for adapting large-scale Visual Language Models (VLMs). However, current MPL methods are fundamentally limited by their optimization of a single, static point representation. This paradigm is inherently brittle, leads to overfitting on base classes, and generalizes poorly to novel or ambiguous categories. We challenge this point paradigm, proposing that robust generalization requires learning a semantic cloud (i.e., a distribution over the embedding space). To achieve this, we introduce Points-to-Clouds (P2C), a novel framework inspired by diffusion models that reframes prompt learning as a dynamic denoising task. At the core of P2C is a dual denoising mechanism: a Dynamic Prompt Denoising (DPD) mechanism perturbs text prompts with sophisticated, annealed noise to learn a smoother semantic landscape, while an auxiliary V-L Mapper denoising loss re-tasks the mapper as a denoising autoencoder. This forces the mapper to reconstruct clean visual prompts from noisy text inputs, ensuring robust cross-modal alignment. Extensive experiments across 11 datasets demonstrate that P2C consistently outperforms strong baselines. On the base-to-novel generalization benchmark, our method achieves a Harmonic Mean of 79.7%, representing a relative improvement of 1.4% over the baseline. The code and models are available at https://vranlee.github.io/P2C/.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22897
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Points to Clouds: Learning Robust Semantic Distributions for Multi-modal Prompts
Li, Weiran
Liu, Yeqiang
Wei, Yijie
Han, Mina
Liu, Xin
Li, Zhenbo
Computer Vision and Pattern Recognition
Multimodal Prompt Learning (MPL) has emerged as a pivotal technique for adapting large-scale Visual Language Models (VLMs). However, current MPL methods are fundamentally limited by their optimization of a single, static point representation. This paradigm is inherently brittle, leads to overfitting on base classes, and generalizes poorly to novel or ambiguous categories. We challenge this point paradigm, proposing that robust generalization requires learning a semantic cloud (i.e., a distribution over the embedding space). To achieve this, we introduce Points-to-Clouds (P2C), a novel framework inspired by diffusion models that reframes prompt learning as a dynamic denoising task. At the core of P2C is a dual denoising mechanism: a Dynamic Prompt Denoising (DPD) mechanism perturbs text prompts with sophisticated, annealed noise to learn a smoother semantic landscape, while an auxiliary V-L Mapper denoising loss re-tasks the mapper as a denoising autoencoder. This forces the mapper to reconstruct clean visual prompts from noisy text inputs, ensuring robust cross-modal alignment. Extensive experiments across 11 datasets demonstrate that P2C consistently outperforms strong baselines. On the base-to-novel generalization benchmark, our method achieves a Harmonic Mean of 79.7%, representing a relative improvement of 1.4% over the baseline. The code and models are available at https://vranlee.github.io/P2C/.
title From Points to Clouds: Learning Robust Semantic Distributions for Multi-modal Prompts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.22897