SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Miller, Kevin, Mishra, Samarth, Gangrade, Aditya, Saenko, Kate, Saligrama, Venkatesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
by: Mishra, Samarth, et al.
Published: (2025)
by: Mishra, Samarth, et al.
Published: (2025)
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
by: Mishra, Samarth, et al.
Published: (2023)
by: Mishra, Samarth, et al.
Published: (2023)
Constrained Linear Thompson Sampling
by: Gangrade, Aditya, et al.
Published: (2025)
by: Gangrade, Aditya, et al.
Published: (2025)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
by: Liu, Aoming, et al.
Published: (2025)
by: Liu, Aoming, et al.
Published: (2025)
Safe Linear Bandits over Unknown Polytopes
by: Gangrade, Aditya, et al.
Published: (2022)
by: Gangrade, Aditya, et al.
Published: (2022)
Data Deletion Can Help in Adaptive RL
by: Budhraja, Param, et al.
Published: (2026)
by: Budhraja, Param, et al.
Published: (2026)
Deep Companion Learning: Enhancing Generalization Through Historical Consistency
by: Zhu, Ruizhao, et al.
Published: (2024)
by: Zhu, Ruizhao, et al.
Published: (2024)
Web Artifact Attacks Disrupt Vision Language Models
by: Qraitem, Maan, et al.
Published: (2025)
by: Qraitem, Maan, et al.
Published: (2025)
Testing the Feasibility of Linear Programs with Bandit Feedback
by: Gangrade, Aditya, et al.
Published: (2024)
by: Gangrade, Aditya, et al.
Published: (2024)
Prompt-Induced Score Variance in Zero-Shot Binary Vision-Language Safety Classification
by: Weng, Charles, et al.
Published: (2026)
by: Weng, Charles, et al.
Published: (2026)
Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting
by: Zhu, Xingyu, et al.
Published: (2024)
by: Zhu, Xingyu, et al.
Published: (2024)
Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
by: Yin, Xiaojie, et al.
Published: (2025)
by: Yin, Xiaojie, et al.
Published: (2025)
Linear Transformers Implicitly Discover Unified Numerical Algorithms
by: Lutz, Patrick, et al.
Published: (2025)
by: Lutz, Patrick, et al.
Published: (2025)
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
by: Petsiuk, Vitali, et al.
Published: (2024)
by: Petsiuk, Vitali, et al.
Published: (2024)
CLAMP: Contrastive LAnguage Model Prompt-tuning
by: Teterwak, Piotr, et al.
Published: (2023)
by: Teterwak, Piotr, et al.
Published: (2023)
Investigating Zero-Shot Diagnostic Pathology in Vision-Language Models with Efficient Prompt Design
by: Sharma, Vasudev, et al.
Published: (2025)
by: Sharma, Vasudev, et al.
Published: (2025)
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
by: Qraitem, Maan, et al.
Published: (2023)
by: Qraitem, Maan, et al.
Published: (2023)
Novel Semantic Prompting for Zero-Shot Action Recognition
by: Iqbal, Salman, et al.
Published: (2026)
by: Iqbal, Salman, et al.
Published: (2026)
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
by: Zhang, Pu, et al.
Published: (2025)
by: Zhang, Pu, et al.
Published: (2025)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
by: Xu, Zhenlin, et al.
Published: (2023)
by: Xu, Zhenlin, et al.
Published: (2023)
Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models
by: Ranasinghe, Yasiru, et al.
Published: (2025)
by: Ranasinghe, Yasiru, et al.
Published: (2025)
Towards Zero-Shot Differential Morphing Attack Detection with Multimodal Large Language Models
by: Shekhawat, Ria, et al.
Published: (2025)
by: Shekhawat, Ria, et al.
Published: (2025)
Zero-Shot Product Attribute Labeling with Vision-Language Models: A Three-Tier Evaluation Framework
by: Shukla, Shubham, et al.
Published: (2026)
by: Shukla, Shubham, et al.
Published: (2026)
Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization
by: Wu, Yuanli, et al.
Published: (2025)
by: Wu, Yuanli, et al.
Published: (2025)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
by: Sharma, Saurav, et al.
Published: (2025)
by: Sharma, Saurav, et al.
Published: (2025)
Efficient and Context-Aware Label Propagation for Zero-/Few-Shot Training-Free Adaptation of Vision-Language Model
by: Li, Yushu, et al.
Published: (2024)
by: Li, Yushu, et al.
Published: (2024)
Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models
by: Choi, In Chong, et al.
Published: (2026)
by: Choi, In Chong, et al.
Published: (2026)
Visual Adaptive Prompting for Compositional Zero-Shot Learning
by: Stein, Kyle, et al.
Published: (2025)
by: Stein, Kyle, et al.
Published: (2025)
Noise is an Efficient Learner for Zero-Shot Vision-Language Models
by: Imam, Raza, et al.
Published: (2025)
by: Imam, Raza, et al.
Published: (2025)
HeatPrompt: Zero-Shot Vision-Language Modeling of Urban Heat Demand from Satellite Images
by: Thota, Kundan, et al.
Published: (2026)
by: Thota, Kundan, et al.
Published: (2026)
Pedestrian Attribute Recognition via CLIP based Prompt Vision-Language Fusion
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
Enhancing Zero-Shot Image Recognition in Vision-Language Models through Human-like Concept Guidance
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
SalientFusion: Context-Aware Compositional Zero-Shot Food Recognition
by: Song, Jiajun, et al.
Published: (2025)
by: Song, Jiajun, et al.
Published: (2025)
Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification
by: Lutz, Patrick, et al.
Published: (2026)
by: Lutz, Patrick, et al.
Published: (2026)
Prompting Language-Informed Distribution for Compositional Zero-Shot Learning
by: Bao, Wentao, et al.
Published: (2023)
by: Bao, Wentao, et al.
Published: (2023)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
by: Chen, Shiming, et al.
Published: (2025)
by: Chen, Shiming, et al.
Published: (2025)
Evaluating Cell Type Inference in Vision Language Models Under Varying Visual Context
by: Singhal, Samarth, et al.
Published: (2025)
by: Singhal, Samarth, et al.
Published: (2025)
Prompts to Summaries: Zero-Shot Language-Guided Video Summarization with Large Language and Video Models
by: Barbara, Mario, et al.
Published: (2025)
by: Barbara, Mario, et al.
Published: (2025)
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
NLPrompt: Noise-Label Prompt Learning for Vision-Language Models
by: Pan, Bikang, et al.
Published: (2024)
by: Pan, Bikang, et al.
Published: (2024)
Similar Items
-
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
by: Mishra, Samarth, et al.
Published: (2025) -
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
by: Mishra, Samarth, et al.
Published: (2023) -
Constrained Linear Thompson Sampling
by: Gangrade, Aditya, et al.
Published: (2025) -
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
by: Liu, Aoming, et al.
Published: (2025) -
Safe Linear Bandits over Unknown Polytopes
by: Gangrade, Aditya, et al.
Published: (2022)