OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: C, Mohamad Hassan N, Gupta, Divyam, Singha, Mainak, Rongali, Sai Bhargav, Jha, Ankit, Khan, Muhammad Haris, Banerjee, Biplab
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917963327602688
author C, Mohamad Hassan N
Gupta, Divyam
Singha, Mainak
Rongali, Sai Bhargav
Jha, Ankit
Khan, Muhammad Haris
Banerjee, Biplab
author_facet C, Mohamad Hassan N
Gupta, Divyam
Singha, Mainak
Rongali, Sai Bhargav
Jha, Ankit
Khan, Muhammad Haris
Banerjee, Biplab
contents We introduce Low-Shot Open-Set Domain Generalization (LSOSDG), a novel paradigm unifying low-shot learning with open-set domain generalization (ODG). While prompt-based methods using models like CLIP have advanced DG, they falter in low-data regimes (e.g., 1-shot) and lack precision in detecting open-set samples with fine-grained semantics related to training classes. To address these challenges, we propose OSLOPROMPT, an advanced prompt-learning framework for CLIP with two core innovations. First, to manage limited supervision across source domains and improve DG, we introduce a domain-agnostic prompt-learning mechanism that integrates adaptable domain-specific cues and visually guided semantic attributes through a novel cross-attention module, besides being supported by learnable domain- and class-generic visual prompts to enhance cross-modal adaptability. Second, to improve outlier rejection during inference, we classify unfamiliar samples as "unknown" and train specialized prompts with systematically synthesized pseudo-open samples that maintain fine-grained relationships to known classes, generated through a targeted query strategy with off-the-shelf foundation models. This strategy enhances feature learning, enabling our model to detect open samples with varied granularity more effectively. Extensive evaluations across five benchmarks demonstrate that OSLOPROMPT establishes a new state-of-the-art in LSOSDG, significantly outperforming existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16106
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP
C, Mohamad Hassan N
Gupta, Divyam
Singha, Mainak
Rongali, Sai Bhargav
Jha, Ankit
Khan, Muhammad Haris
Banerjee, Biplab
Computer Vision and Pattern Recognition
We introduce Low-Shot Open-Set Domain Generalization (LSOSDG), a novel paradigm unifying low-shot learning with open-set domain generalization (ODG). While prompt-based methods using models like CLIP have advanced DG, they falter in low-data regimes (e.g., 1-shot) and lack precision in detecting open-set samples with fine-grained semantics related to training classes. To address these challenges, we propose OSLOPROMPT, an advanced prompt-learning framework for CLIP with two core innovations. First, to manage limited supervision across source domains and improve DG, we introduce a domain-agnostic prompt-learning mechanism that integrates adaptable domain-specific cues and visually guided semantic attributes through a novel cross-attention module, besides being supported by learnable domain- and class-generic visual prompts to enhance cross-modal adaptability. Second, to improve outlier rejection during inference, we classify unfamiliar samples as "unknown" and train specialized prompts with systematically synthesized pseudo-open samples that maintain fine-grained relationships to known classes, generated through a targeted query strategy with off-the-shelf foundation models. This strategy enhances feature learning, enabling our model to detect open samples with varied granularity more effectively. Extensive evaluations across five benchmarks demonstrate that OSLOPROMPT establishes a new state-of-the-art in LSOSDG, significantly outperforming existing methods.
title OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.16106