Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions
Fuente:
arXiv
Saved in:
| Main Authors: | Saha, Oindrila, Van Horn, Grant, Maji, Subhransu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generate, Transduct, Adapt: Iterative Transduction with VLMs
by: Saha, Oindrila, et al.
Published: (2025)
by: Saha, Oindrila, et al.
Published: (2025)
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
by: Lawrence, Logan, et al.
Published: (2025)
by: Lawrence, Logan, et al.
Published: (2025)
Not All Birds Look The Same: Identity-Preserving Generation For Birds
by: Sun, Aaron, et al.
Published: (2025)
by: Sun, Aaron, et al.
Published: (2025)
Merlin L48 Spectrogram Dataset
by: Sun, Aaron, et al.
Published: (2025)
by: Sun, Aaron, et al.
Published: (2025)
3D Space as a Scratchpad for Editable Text-to-Image Generation
by: Saha, Oindrila, et al.
Published: (2026)
by: Saha, Oindrila, et al.
Published: (2026)
Human-in-the-Loop Visual Re-ID for Population Size Estimation
by: Perez, Gustavo, et al.
Published: (2023)
by: Perez, Gustavo, et al.
Published: (2023)
Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study
by: Liu, Wuao, et al.
Published: (2026)
by: Liu, Wuao, et al.
Published: (2026)
SIGMA-GEN: Structure and Identity Guided Multi-subject Assembly for Image Generation
by: Saha, Oindrila, et al.
Published: (2025)
by: Saha, Oindrila, et al.
Published: (2025)
RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
by: Lawrence, Logan, et al.
Published: (2026)
by: Lawrence, Logan, et al.
Published: (2026)
Consensus-Driven Active Model Selection
by: Kay, Justin, et al.
Published: (2025)
by: Kay, Justin, et al.
Published: (2025)
Moment Sampling in Video LLMs for Long-Form Video QA
by: Chasmai, Mustafa, et al.
Published: (2025)
by: Chasmai, Mustafa, et al.
Published: (2025)
WildSAT: Learning Satellite Image Representations from Wildlife Observations
by: Daroya, Rangel, et al.
Published: (2024)
by: Daroya, Rangel, et al.
Published: (2024)
YouDream: Generating Anatomically Controllable Consistent Text-to-3D Animals
by: Mishra, Sandeep, et al.
Published: (2024)
by: Mishra, Sandeep, et al.
Published: (2024)
Task2Box: Box Embeddings for Modeling Asymmetric Task Relationships
by: Daroya, Rangel, et al.
Published: (2024)
by: Daroya, Rangel, et al.
Published: (2024)
Active Measurement of Two-Point Correlations
by: Hamilton, Max, et al.
Published: (2026)
by: Hamilton, Max, et al.
Published: (2026)
Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks
by: Khan, Muhammad Saif Ullah, et al.
Published: (2024)
by: Khan, Muhammad Saif Ullah, et al.
Published: (2024)
BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs
by: Baranwal, Aaditya, et al.
Published: (2026)
by: Baranwal, Aaditya, et al.
Published: (2026)
Feedforward Few-shot Species Range Estimation
by: Lange, Christian, et al.
Published: (2025)
by: Lange, Christian, et al.
Published: (2025)
C3DAG: Controlled 3D Animal Generation using 3D pose guidance
by: Mishra, Sandeep, et al.
Published: (2024)
by: Mishra, Sandeep, et al.
Published: (2024)
VIDMP3: Video Editing by Representing Motion with Pose and Position Priors
by: Mishra, Sandeep, et al.
Published: (2025)
by: Mishra, Sandeep, et al.
Published: (2025)
Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning
by: Wu, Xiangyu, et al.
Published: (2025)
by: Wu, Xiangyu, et al.
Published: (2025)
Liberating Seen Classes: Boosting Few-Shot and Zero-Shot Text Classification via Anchor Generation and Classification Reframing
by: Liu, Han, et al.
Published: (2024)
by: Liu, Han, et al.
Published: (2024)
Will It Zero-Shot?: Predicting Zero-Shot Classification Performance For Arbitrary Queries
by: Robbins, Kevin, et al.
Published: (2026)
by: Robbins, Kevin, et al.
Published: (2026)
Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers
by: Eltahir, Mohamed, et al.
Published: (2025)
by: Eltahir, Mohamed, et al.
Published: (2025)
Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs
by: Hasanebrahimi, Afsaneh, et al.
Published: (2026)
by: Hasanebrahimi, Afsaneh, et al.
Published: (2026)
Exploring Prompt Alignment with Clinical Factors in Zero-Shot Segmentation VLMs for NSCLC Tumor Segmentation
by: Pai, Suraj, et al.
Published: (2026)
by: Pai, Suraj, et al.
Published: (2026)
Improving Satellite Imagery Masking using Multi-task and Transfer Learning
by: Daroya, Rangel, et al.
Published: (2024)
by: Daroya, Rangel, et al.
Published: (2024)
Zero-Shot Aerial Object Detection with Visual Description Regularization
by: Zang, Zhengqing, et al.
Published: (2024)
by: Zang, Zhengqing, et al.
Published: (2024)
Active Measurement: Efficient Estimation at Scale
by: Hamilton, Max, et al.
Published: (2025)
by: Hamilton, Max, et al.
Published: (2025)
Text-guided Zero-Shot Object Localization
by: Wang, Jingjing, et al.
Published: (2024)
by: Wang, Jingjing, et al.
Published: (2024)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
by: Luo, Kun, et al.
Published: (2026)
by: Luo, Kun, et al.
Published: (2026)
Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment
by: Gan, Lubin, et al.
Published: (2025)
by: Gan, Lubin, et al.
Published: (2025)
AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection
by: Cao, Yunkang, et al.
Published: (2024)
by: Cao, Yunkang, et al.
Published: (2024)
ChartZero: Synthetic Priors Enable Zero Shot Chart Data Extraction
by: Islam, Md Touhidul, et al.
Published: (2026)
by: Islam, Md Touhidul, et al.
Published: (2026)
AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
by: Jiang, Yichen, et al.
Published: (2025)
by: Jiang, Yichen, et al.
Published: (2025)
Learning to Adapt Category Consistent Meta-Feature of CLIP for Few-Shot Classification
by: Shi, Jiaying, et al.
Published: (2024)
by: Shi, Jiaying, et al.
Published: (2024)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024)
by: Romero, David, et al.
Published: (2024)
LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation
by: Li, Jiachen, et al.
Published: (2025)
by: Li, Jiachen, et al.
Published: (2025)
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
by: Motamed, Saman, et al.
Published: (2024)
by: Motamed, Saman, et al.
Published: (2024)
Similar Items
-
Generate, Transduct, Adapt: Iterative Transduction with VLMs
by: Saha, Oindrila, et al.
Published: (2025) -
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
by: Lawrence, Logan, et al.
Published: (2025) -
Not All Birds Look The Same: Identity-Preserving Generation For Birds
by: Sun, Aaron, et al.
Published: (2025) -
Merlin L48 Spectrogram Dataset
by: Sun, Aaron, et al.
Published: (2025) -
3D Space as a Scratchpad for Editable Text-to-Image Generation
by: Saha, Oindrila, et al.
Published: (2026)