What Do You See? Enhancing Zero-Shot Image Classification with Multimodal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Abdelhamed, Abdelrahman, Afifi, Mahmoud, Go, Alec |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
by: De Nadai, Marco, et al.
Published: (2025)
by: De Nadai, Marco, et al.
Published: (2025)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026)
by: Song, Junha, et al.
Published: (2026)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
by: Atabuzzaman, Md., et al.
Published: (2025)
by: Atabuzzaman, Md., et al.
Published: (2025)
ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models
by: Gong, Bingchen, et al.
Published: (2024)
by: Gong, Bingchen, et al.
Published: (2024)
Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models
by: Scius-Bertrand, Anna, et al.
Published: (2024)
by: Scius-Bertrand, Anna, et al.
Published: (2024)
Optimizing Illuminant Estimation in Dual-Exposure HDR Imaging
by: Afifi, Mahmoud, et al.
Published: (2024)
by: Afifi, Mahmoud, et al.
Published: (2024)
What You See is (Usually) What You Get: Multimodal Prototype Networks that Abstain from Expensive Modalities
by: Bahng, Muchang, et al.
Published: (2025)
by: Bahng, Muchang, et al.
Published: (2025)
Raw-JPEG Adapter: Efficient Raw Image Compression with JPEG
by: Afifi, Mahmoud, et al.
Published: (2025)
by: Afifi, Mahmoud, et al.
Published: (2025)
Enhancing Remote Sensing Vision-Language Models for Zero-Shot Scene Classification
by: Khoury, Karim El, et al.
Published: (2024)
by: Khoury, Karim El, et al.
Published: (2024)
Improved Mapping Between Illuminations and Sensors for RAW Images
by: Punnappurath, Abhijith, et al.
Published: (2025)
by: Punnappurath, Abhijith, et al.
Published: (2025)
What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze Estimation
by: Cheng, Yihua, et al.
Published: (2024)
by: Cheng, Yihua, et al.
Published: (2024)
Zero-Shot Anomaly Detection in Battery Thermal Images Using Visual Question Answering with Prior Knowledge
by: Astrid, Marcella, et al.
Published: (2025)
by: Astrid, Marcella, et al.
Published: (2025)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
by: Mitra, Chancharik, et al.
Published: (2024)
by: Mitra, Chancharik, et al.
Published: (2024)
Modular Neural Image Signal Processing
by: Afifi, Mahmoud, et al.
Published: (2025)
by: Afifi, Mahmoud, et al.
Published: (2025)
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
by: Deng, Huilin, et al.
Published: (2024)
by: Deng, Huilin, et al.
Published: (2024)
Towards Zero-Shot Differential Morphing Attack Detection with Multimodal Large Language Models
by: Shekhawat, Ria, et al.
Published: (2025)
by: Shekhawat, Ria, et al.
Published: (2025)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
by: Elhenawy, Mohammed, et al.
Published: (2025)
by: Elhenawy, Mohammed, et al.
Published: (2025)
Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models
by: Xu, Jiacong, et al.
Published: (2025)
by: Xu, Jiacong, et al.
Published: (2025)
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
by: Hu, Jinyi, et al.
Published: (2023)
by: Hu, Jinyi, et al.
Published: (2023)
Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs
by: Kanade, Aditya, et al.
Published: (2025)
by: Kanade, Aditya, et al.
Published: (2025)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
by: Choi, Yura, et al.
Published: (2026)
by: Choi, Yura, et al.
Published: (2026)
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
by: Bora, Maheswar, et al.
Published: (2025)
by: Bora, Maheswar, et al.
Published: (2025)
Pushing Boundaries: Exploring Zero Shot Object Classification with Large Multimodal Models
by: Islam, Ashhadul, et al.
Published: (2023)
by: Islam, Ashhadul, et al.
Published: (2023)
Large Language Models are Good Prompt Learners for Low-Shot Image Classification
by: Zheng, Zhaoheng, et al.
Published: (2023)
by: Zheng, Zhaoheng, et al.
Published: (2023)
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
by: Zhang, Jialiang, et al.
Published: (2026)
by: Zhang, Jialiang, et al.
Published: (2026)
Seeing Beyond Classes: Zero-Shot Grounded Situation Recognition via Language Explainer
by: Lei, Jiaming, et al.
Published: (2024)
by: Lei, Jiaming, et al.
Published: (2024)
Learning Camera-Agnostic White-Balance Preferences
by: Zhao, Luxi, et al.
Published: (2025)
by: Zhao, Luxi, et al.
Published: (2025)
What Do You See in Common? Learning Hierarchical Prototypes over Tree-of-Life to Discover Evolutionary Traits
by: Manogaran, Harish Babu, et al.
Published: (2024)
by: Manogaran, Harish Babu, et al.
Published: (2024)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
by: He, Zhentao, et al.
Published: (2025)
by: He, Zhentao, et al.
Published: (2025)
How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR Heads
by: Baek, Ingeol, et al.
Published: (2025)
by: Baek, Ingeol, et al.
Published: (2025)
MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?
by: Dai, Shiqi, et al.
Published: (2025)
by: Dai, Shiqi, et al.
Published: (2025)
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025)
by: Kala, Divy, et al.
Published: (2025)
Seeing the Arrow of Time in Large Multimodal Models
by: Xue, Zihui, et al.
Published: (2025)
by: Xue, Zihui, et al.
Published: (2025)
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons
by: Xu, Zhe, et al.
Published: (2024)
by: Xu, Zhe, et al.
Published: (2024)
Will It Zero-Shot?: Predicting Zero-Shot Classification Performance For Arbitrary Queries
by: Robbins, Kevin, et al.
Published: (2026)
by: Robbins, Kevin, et al.
Published: (2026)
ZeroSlide: Is Zero-Shot Classification Adequate for Lifelong Learning in Whole-Slide Image Analysis in the Era of Pathology Vision-Language Foundation Models?
by: Bui, Doanh C., et al.
Published: (2025)
by: Bui, Doanh C., et al.
Published: (2025)
Seeing the Unseen: Towards Zero-Shot Inspection for Wind Turbine Blades using Knowledge-Augmented Vision Language Models
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Zero-Shot Video Semantic Segmentation based on Pre-Trained Diffusion Models
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
MADS: Multi-Attribute Document Supervision for Zero-Shot Image Classification
by: Qu, Xiangyan, et al.
Published: (2025)
by: Qu, Xiangyan, et al.
Published: (2025)
Similar Items
-
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
by: De Nadai, Marco, et al.
Published: (2025) -
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026) -
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025) -
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
by: Atabuzzaman, Md., et al.
Published: (2025) -
ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models
by: Gong, Bingchen, et al.
Published: (2024)