On Large Multimodal Models as Open-World Image Classifiers
Fuente:
arXiv
Saved in:
| Main Authors: | Conti, Alessandro, Mancini, Massimiliano, Fini, Enrico, Wang, Yiming, Rota, Paolo, Ricci, Elisa |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vocabulary-free Image Classification
by: Conti, Alessandro, et al.
Published: (2023)
by: Conti, Alessandro, et al.
Published: (2023)
Vocabulary-free Image Classification and Semantic Segmentation
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
Large Multimodal Models as General In-Context Classifiers
by: Garosi, Marco, et al.
Published: (2026)
by: Garosi, Marco, et al.
Published: (2026)
Test-Time Zero-Shot Temporal Action Localization
by: Liberatori, Benedetta, et al.
Published: (2024)
by: Liberatori, Benedetta, et al.
Published: (2024)
Compositional Caching for Training-free Open-vocabulary Attribute Detection
by: Garosi, Marco, et al.
Published: (2025)
by: Garosi, Marco, et al.
Published: (2025)
ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
by: Liberatori, Benedetta, et al.
Published: (2025)
by: Liberatori, Benedetta, et al.
Published: (2025)
Zero-Shot Temporal Action Localization Through Textual Guidance
by: Liberatori, Benedetta, et al.
Published: (2026)
by: Liberatori, Benedetta, et al.
Published: (2026)
Harnessing Large Language Models for Training-free Video Anomaly Detection
by: Zanella, Luca, et al.
Published: (2024)
by: Zanella, Luca, et al.
Published: (2024)
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
by: De Min, Thomas, et al.
Published: (2026)
by: De Min, Thomas, et al.
Published: (2026)
Retrieval-enriched zero-shot image classification in low-resource domains
by: Dall'Asen, Nicola, et al.
Published: (2024)
by: Dall'Asen, Nicola, et al.
Published: (2024)
Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifiers
by: Guimard, Quentin, et al.
Published: (2025)
by: Guimard, Quentin, et al.
Published: (2025)
Training-Free Personalization via Retrieval and Reasoning on Fingerprints
by: Das, Deepayan, et al.
Published: (2025)
by: Das, Deepayan, et al.
Published: (2025)
Training-free Online Video Step Grounding
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering
by: Das, Deepayan, et al.
Published: (2024)
by: Das, Deepayan, et al.
Published: (2024)
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
by: Caldarella, Simone, et al.
Published: (2024)
by: Caldarella, Simone, et al.
Published: (2024)
Specificity-aware reinforcement learning for fine-grained open-world classification
by: Angheben, Samuele, et al.
Published: (2026)
by: Angheben, Samuele, et al.
Published: (2026)
Can Text-to-Video Generation help Video-Language Alignment?
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
Unlearning Personal Data from a Single Image
by: De Min, Thomas, et al.
Published: (2024)
by: De Min, Thomas, et al.
Published: (2024)
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
by: Farina, Matteo, et al.
Published: (2025)
by: Farina, Matteo, et al.
Published: (2025)
Frustratingly Easy Test-Time Adaptation of Vision-Language Models
by: Farina, Matteo, et al.
Published: (2024)
by: Farina, Matteo, et al.
Published: (2024)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
by: Berasi, Davide, et al.
Published: (2025)
by: Berasi, Davide, et al.
Published: (2025)
Multimodal Large Language Models as Image Classifiers
by: Kisel, Nikita, et al.
Published: (2026)
by: Kisel, Nikita, et al.
Published: (2026)
MULTIFLOW: Shifting Towards Task-Agnostic Vision-Language Pruning
by: Farina, Matteo, et al.
Published: (2024)
by: Farina, Matteo, et al.
Published: (2024)
Text-Enhanced Zero-Shot Action Recognition: A training-free approach
by: Bosetti, Massimo, et al.
Published: (2024)
by: Bosetti, Massimo, et al.
Published: (2024)
Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection
by: Tonini, Francesco, et al.
Published: (2025)
by: Tonini, Francesco, et al.
Published: (2025)
Towards Unconstrained Human-Object Interaction
by: Tonini, Francesco, et al.
Published: (2026)
by: Tonini, Francesco, et al.
Published: (2026)
Test-time Vocabulary Adaptation for Language-driven Object Detection
by: Liu, Mingxuan, et al.
Published: (2025)
by: Liu, Mingxuan, et al.
Published: (2025)
From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition
by: Gentile, Francesco, et al.
Published: (2026)
by: Gentile, Francesco, et al.
Published: (2026)
Less is more: Summarizing Patch Tokens for efficient Multi-Label Class-Incremental Learning
by: De Min, Thomas, et al.
Published: (2024)
by: De Min, Thomas, et al.
Published: (2024)
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
by: Huang, Yiran, et al.
Published: (2025)
by: Huang, Yiran, et al.
Published: (2025)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
by: Guimard, Quentin, et al.
Published: (2026)
by: Guimard, Quentin, et al.
Published: (2026)
OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
by: Zhang, Zhenhao, et al.
Published: (2025)
by: Zhang, Zhenhao, et al.
Published: (2025)
OpenBias: Open-set Bias Detection in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024)
by: D'Incà, Moreno, et al.
Published: (2024)
Scaling Laws for Native Multimodal Models
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
Vision-by-Language for Training-Free Compositional Image Retrieval
by: Karthik, Shyamgopal, et al.
Published: (2023)
by: Karthik, Shyamgopal, et al.
Published: (2023)
Classifier-Centric Adaptive Framework for Open-Vocabulary Camouflaged Object Segmentation
by: Zhang, Hanyu, et al.
Published: (2025)
by: Zhang, Hanyu, et al.
Published: (2025)
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
by: Tur, Anil Osman, et al.
Published: (2024)
by: Tur, Anil Osman, et al.
Published: (2024)
Novel class discovery meets foundation models for 3D semantic segmentation
by: Riz, Luigi, et al.
Published: (2023)
by: Riz, Luigi, et al.
Published: (2023)
Similar Items
-
Vocabulary-free Image Classification
by: Conti, Alessandro, et al.
Published: (2023) -
Vocabulary-free Image Classification and Semantic Segmentation
by: Conti, Alessandro, et al.
Published: (2024) -
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024) -
Large Multimodal Models as General In-Context Classifiers
by: Garosi, Marco, et al.
Published: (2026) -
Test-Time Zero-Shot Temporal Action Localization
by: Liberatori, Benedetta, et al.
Published: (2024)