Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Yunqi, An, Sohyun, Bai, Andrew, Lin, Neil Y. C., Hsieh, Cho-Jui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Reward Hacking in Text-to-Image Reinforcement Learning
by: Hong, Yunqi, et al.
Published: (2026)
by: Hong, Yunqi, et al.
Published: (2026)
IRIS: Intrinsic Reward Image Synthesis
by: Chen, Yihang, et al.
Published: (2025)
by: Chen, Yihang, et al.
Published: (2025)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
by: Zhou, Yu, et al.
Published: (2025)
by: Zhou, Yu, et al.
Published: (2025)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
by: Li, Juncheng, et al.
Published: (2023)
by: Li, Juncheng, et al.
Published: (2023)
QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
by: Kao, Kuei-Chun, et al.
Published: (2025)
by: Kao, Kuei-Chun, et al.
Published: (2025)
Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
by: Hong, Yunqi, et al.
Published: (2025)
by: Hong, Yunqi, et al.
Published: (2025)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
by: Atabuzzaman, Md., et al.
Published: (2025)
by: Atabuzzaman, Md., et al.
Published: (2025)
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
by: Bai, Andrew, et al.
Published: (2025)
by: Bai, Andrew, et al.
Published: (2025)
Embedding Space Selection for Detecting Memorization and Fingerprinting in Generative Models
by: He, Jack, et al.
Published: (2024)
by: He, Jack, et al.
Published: (2024)
Uncertainty-Guided Selective Adaptation Enables Cross-Platform Predictive Fluorescence Microscopy
by: Yang, Kai-Wen K., et al.
Published: (2025)
by: Yang, Kai-Wen K., et al.
Published: (2025)
Attributed Synthetic Data Generation for Zero-shot Domain-specific Image Classification
by: Wang, Shijian, et al.
Published: (2025)
by: Wang, Shijian, et al.
Published: (2025)
SPECIAL: Zero-shot Hyperspectral Image Classification With CLIP
by: Pang, Li, et al.
Published: (2025)
by: Pang, Li, et al.
Published: (2025)
MuSc: Zero-Shot Industrial Anomaly Classification and Segmentation with Mutual Scoring of the Unlabeled Images
by: Li, Xurui, et al.
Published: (2024)
by: Li, Xurui, et al.
Published: (2024)
MuLan: Multimodal-LLM Agent for Progressive and Interactive Multi-Object Diffusion
by: Li, Sen, et al.
Published: (2024)
by: Li, Sen, et al.
Published: (2024)
ToolFG: Towards Well-Grounded Fine-Grained Image Classification
by: Xue, Yu, et al.
Published: (2026)
by: Xue, Yu, et al.
Published: (2026)
MuSc-V2: Zero-Shot Multimodal Industrial Anomaly Classification and Segmentation with Mutual Scoring of Unlabeled Samples
by: Li, Xurui, et al.
Published: (2025)
by: Li, Xurui, et al.
Published: (2025)
Diverse and Tailored Image Generation for Zero-shot Multi-label Classification
by: Zhang, Kaixin, et al.
Published: (2024)
by: Zhang, Kaixin, et al.
Published: (2024)
FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation
by: Shao, Dian, et al.
Published: (2026)
by: Shao, Dian, et al.
Published: (2026)
Multi-method Integration with Confidence-based Weighting for Zero-shot Image Classification
by: Yin, Siqi, et al.
Published: (2024)
by: Yin, Siqi, et al.
Published: (2024)
Data-Efficient Generalization for Zero-shot Composed Image Retrieval
by: Chen, Zining, et al.
Published: (2025)
by: Chen, Zining, et al.
Published: (2025)
MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
by: Du, Yipeng, et al.
Published: (2025)
by: Du, Yipeng, et al.
Published: (2025)
Image to Pseudo-Episode: Boosting Few-Shot Segmentation by Unlabeled Data
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
ZeroStereo: Zero-shot Stereo Matching from Single Images
by: Wang, Xianqi, et al.
Published: (2025)
by: Wang, Xianqi, et al.
Published: (2025)
HMIL: Hierarchical Multi-Instance Learning for Fine-Grained Whole Slide Image Classification
by: Jin, Cheng, et al.
Published: (2024)
by: Jin, Cheng, et al.
Published: (2024)
Fine-Grained Scene Image Classification with Modality-Agnostic Adapter
by: Wang, Yiqun, et al.
Published: (2024)
by: Wang, Yiqun, et al.
Published: (2024)
Zero-shot Shape Classification of Nanoparticles in SEM Images using Vision Foundation Models
by: Barnatan, Freida, et al.
Published: (2025)
by: Barnatan, Freida, et al.
Published: (2025)
Zero-shot Quantization: A Comprehensive Survey
by: Kim, Minjun, et al.
Published: (2025)
by: Kim, Minjun, et al.
Published: (2025)
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
by: Tur, Anil Osman, et al.
Published: (2024)
by: Tur, Anil Osman, et al.
Published: (2024)
MCFNet: A Multimodal Collaborative Fusion Network for Fine-Grained Semantic Classification
by: Qiao, Yang, et al.
Published: (2025)
by: Qiao, Yang, et al.
Published: (2025)
Fine-Grained ImageNet Classification in the Wild
by: Lymperaiou, Maria, et al.
Published: (2023)
by: Lymperaiou, Maria, et al.
Published: (2023)
Zero-Shot Head Swapping in Real-World Scenarios
by: Kang, Taewoong, et al.
Published: (2025)
by: Kang, Taewoong, et al.
Published: (2025)
TIP: Tabular-Image Pre-training for Multimodal Classification with Incomplete Data
by: Du, Siyi, et al.
Published: (2024)
by: Du, Siyi, et al.
Published: (2024)
Unlock the Power of Unlabeled Data in Language Driving Model
by: Wang, Chaoqun, et al.
Published: (2025)
by: Wang, Chaoqun, et al.
Published: (2025)
Fine-gained Zero-shot Video Sampling
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
JIGMARK: A Black-Box Approach for Enhancing Image Watermarks against Diffusion Model Edits
by: Pan, Minzhou, et al.
Published: (2024)
by: Pan, Minzhou, et al.
Published: (2024)
Zero-shot Image Editing with Reference Imitation
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Zero-shot Composed Text-Image Retrieval
by: Liu, Yikun, et al.
Published: (2023)
by: Liu, Yikun, et al.
Published: (2023)
Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
by: Yang, William, et al.
Published: (2025)
by: Yang, William, et al.
Published: (2025)
Fine-Grained Zero-Shot Object Detection
by: Ma, Hongxu, et al.
Published: (2025)
by: Ma, Hongxu, et al.
Published: (2025)
SOOD++: Leveraging Unlabeled Data to Boost Oriented Object Detection
by: Liang, Dingkang, et al.
Published: (2024)
by: Liang, Dingkang, et al.
Published: (2024)
Similar Items
-
Understanding Reward Hacking in Text-to-Image Reinforcement Learning
by: Hong, Yunqi, et al.
Published: (2026) -
IRIS: Intrinsic Reward Image Synthesis
by: Chen, Yihang, et al.
Published: (2025) -
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
by: Zhou, Yu, et al.
Published: (2025) -
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
by: Li, Juncheng, et al.
Published: (2023) -
QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
by: Kao, Kuei-Chun, et al.
Published: (2025)