Multimodal Large Language Models as Image Classifiers
Fuente:
arXiv
Saved in:
| Main Authors: | Kisel, Nikita, Volkov, Illia, Janouskova, Klara, Matas, Jiri |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Image Recognition with Vision and Language Embeddings of VLMs
by: Volkov, Illia, et al.
Published: (2025)
by: Volkov, Illia, et al.
Published: (2025)
Flaws of ImageNet, Computer Vision's Favourite Dataset
by: Kisel, Nikita, et al.
Published: (2024)
by: Kisel, Nikita, et al.
Published: (2024)
Bringing the Context Back into Object Recognition, Robustly
by: Janouskova, Klara, et al.
Published: (2024)
by: Janouskova, Klara, et al.
Published: (2024)
Robust Context-Aware Object Recognition
by: Janouskova, Klara, et al.
Published: (2025)
by: Janouskova, Klara, et al.
Published: (2025)
Koo-Fu CLIP: Closed-Form Adaptation of Vision-Language Models via Fukunaga-Koontz Linear Discriminant Analysis
by: Suchanek, Matej, et al.
Published: (2026)
by: Suchanek, Matej, et al.
Published: (2026)
Single Image Test-Time Adaptation for Segmentation
by: Janouskova, Klara, et al.
Published: (2023)
by: Janouskova, Klara, et al.
Published: (2023)
FungiTastic: A multi-modal dataset and benchmark for image categorization
by: Picek, Lukas, et al.
Published: (2024)
by: Picek, Lukas, et al.
Published: (2024)
Accurate Planar Tracking With Robust Re-Detection
by: Serych, Jonas, et al.
Published: (2026)
by: Serych, Jonas, et al.
Published: (2026)
Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous Circle
by: Purkrabek, Miroslav, et al.
Published: (2024)
by: Purkrabek, Miroslav, et al.
Published: (2024)
Video shutter angle estimation using optical flow and linear blur
by: Korcak, David, et al.
Published: (2023)
by: Korcak, David, et al.
Published: (2023)
Improving 2D Human Pose Estimation in Rare Camera Views with Synthetic Data
by: Purkrabek, Miroslav, et al.
Published: (2023)
by: Purkrabek, Miroslav, et al.
Published: (2023)
ProbPose: A Probabilistic Approach to 2D Human Pose Estimation
by: Purkrabek, Miroslav, et al.
Published: (2024)
by: Purkrabek, Miroslav, et al.
Published: (2024)
On Large Multimodal Models as Open-World Image Classifiers
by: Conti, Alessandro, et al.
Published: (2025)
by: Conti, Alessandro, et al.
Published: (2025)
SAM2RL: Towards Reinforcement Learning Memory Control in Segment Anything Model 2
by: Adamyan, Alen, et al.
Published: (2025)
by: Adamyan, Alen, et al.
Published: (2025)
Animal Identification with Independent Foreground and Background Modeling
by: Picek, Lukas, et al.
Published: (2024)
by: Picek, Lukas, et al.
Published: (2024)
SAM-pose2seg: Pose-Guided Human Instance Segmentation in Crowds
by: Kolomiiets, Constantin, et al.
Published: (2026)
by: Kolomiiets, Constantin, et al.
Published: (2026)
BBoxMaskPose v2: Expanding Mutual Conditioning to 3D
by: Purkrabek, Miroslav, et al.
Published: (2026)
by: Purkrabek, Miroslav, et al.
Published: (2026)
Dense Matchers for Dense Tracking
by: Jelínek, Tomáš, et al.
Published: (2024)
by: Jelínek, Tomáš, et al.
Published: (2024)
MFTIQ: Multi-Flow Tracker with Independent Matching Quality Estimation
by: Serych, Jonas, et al.
Published: (2024)
by: Serych, Jonas, et al.
Published: (2024)
Human Pose-Constrained UV Map Estimation
by: Suchanek, Matej, et al.
Published: (2025)
by: Suchanek, Matej, et al.
Published: (2025)
Unlocking the Hidden Potential of CLIP in Generalizable Deepfake Detection
by: Yermakov, Andrii, et al.
Published: (2025)
by: Yermakov, Andrii, et al.
Published: (2025)
PixOOD: Pixel-Level Out-of-Distribution Detection
by: Vojíř, Tomáš, et al.
Published: (2024)
by: Vojíř, Tomáš, et al.
Published: (2024)
Large Multimodal Models as General In-Context Classifiers
by: Garosi, Marco, et al.
Published: (2026)
by: Garosi, Marco, et al.
Published: (2026)
HelixTrack: Event-Based Tracking and RPM Estimation of Propeller-like Objects
by: Spetlik, Radim, et al.
Published: (2026)
by: Spetlik, Radim, et al.
Published: (2026)
The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection
by: Yermakov, Andrii, et al.
Published: (2026)
by: Yermakov, Andrii, et al.
Published: (2026)
Global-Aware Edge Prioritization for Pose Graph Initialization
by: Wei, Tong, et al.
Published: (2026)
by: Wei, Tong, et al.
Published: (2026)
Deepfake Detection that Generalizes Across Benchmarks
by: Yermakov, Andrii, et al.
Published: (2025)
by: Yermakov, Andrii, et al.
Published: (2025)
Breaking the Frame: Visual Place Recognition by Overlap Prediction
by: Wei, Tong, et al.
Published: (2024)
by: Wei, Tong, et al.
Published: (2024)
A New Dataset and a Distractor-Aware Architecture for Transparent Object Tracking
by: Lukezic, Alan, et al.
Published: (2024)
by: Lukezic, Alan, et al.
Published: (2024)
Large Language Models for Multimodal Deformable Image Registration
by: Ma, Mingrui, et al.
Published: (2024)
by: Ma, Mingrui, et al.
Published: (2024)
Safety of Multimodal Large Language Models on Images and Texts
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Three Things to Know about Deep Metric Learning
by: Patel, Yash, et al.
Published: (2024)
by: Patel, Yash, et al.
Published: (2024)
DAD-3DHeads: A Large-scale Dense, Accurate and Diverse Dataset for 3D Head Alignment from a Single Image
by: Martyniuk, Tetiana, et al.
Published: (2022)
by: Martyniuk, Tetiana, et al.
Published: (2022)
Are Multimodal Large Language Models Good Annotators for Image Tagging?
by: Xie, Ming-Kun, et al.
Published: (2026)
by: Xie, Ming-Kun, et al.
Published: (2026)
EE3P: Event-based Estimation of Periodic Phenomena Properties
by: Kolář, Jakub, et al.
Published: (2024)
by: Kolář, Jakub, et al.
Published: (2024)
EEPPR: Event-based Estimation of Periodic Phenomena Rate using Correlation in 3D
by: Kolář, Jakub, et al.
Published: (2024)
by: Kolář, Jakub, et al.
Published: (2024)
VIP: Versatile Image Outpainting Empowered by Multimodal Large Language Model
by: Yang, Jinze, et al.
Published: (2024)
by: Yang, Jinze, et al.
Published: (2024)
MIBench: Evaluating Multimodal Large Language Models over Multiple Images
by: Liu, Haowei, et al.
Published: (2024)
by: Liu, Haowei, et al.
Published: (2024)
Point Cloud Color Constancy
by: Xing, Xiaoyan, et al.
Published: (2021)
by: Xing, Xiaoyan, et al.
Published: (2021)
WildFusion: Individual Animal Identification with Calibrated Similarity Fusion
by: Cermak, Vojtěch, et al.
Published: (2024)
by: Cermak, Vojtěch, et al.
Published: (2024)
Similar Items
-
Image Recognition with Vision and Language Embeddings of VLMs
by: Volkov, Illia, et al.
Published: (2025) -
Flaws of ImageNet, Computer Vision's Favourite Dataset
by: Kisel, Nikita, et al.
Published: (2024) -
Bringing the Context Back into Object Recognition, Robustly
by: Janouskova, Klara, et al.
Published: (2024) -
Robust Context-Aware Object Recognition
by: Janouskova, Klara, et al.
Published: (2025) -
Koo-Fu CLIP: Closed-Form Adaptation of Vision-Language Models via Fukunaga-Koontz Linear Discriminant Analysis
by: Suchanek, Matej, et al.
Published: (2026)