Large Multimodal Models as General In-Context Classifiers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Garosi, Marco, Farina, Matteo, Conti, Alessandro, Mancini, Massimiliano, Ricci, Elisa |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Compositional Caching for Training-free Open-vocabulary Attribute Detection
von: Garosi, Marco, et al.
Veröffentlicht: (2025)
von: Garosi, Marco, et al.
Veröffentlicht: (2025)
On Large Multimodal Models as Open-World Image Classifiers
von: Conti, Alessandro, et al.
Veröffentlicht: (2025)
von: Conti, Alessandro, et al.
Veröffentlicht: (2025)
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
von: Farina, Matteo, et al.
Veröffentlicht: (2025)
von: Farina, Matteo, et al.
Veröffentlicht: (2025)
Frustratingly Easy Test-Time Adaptation of Vision-Language Models
von: Farina, Matteo, et al.
Veröffentlicht: (2024)
von: Farina, Matteo, et al.
Veröffentlicht: (2024)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
von: Berasi, Davide, et al.
Veröffentlicht: (2025)
von: Berasi, Davide, et al.
Veröffentlicht: (2025)
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
von: De Min, Thomas, et al.
Veröffentlicht: (2026)
von: De Min, Thomas, et al.
Veröffentlicht: (2026)
MULTIFLOW: Shifting Towards Task-Agnostic Vision-Language Pruning
von: Farina, Matteo, et al.
Veröffentlicht: (2024)
von: Farina, Matteo, et al.
Veröffentlicht: (2024)
Vocabulary-free Image Classification and Semantic Segmentation
von: Conti, Alessandro, et al.
Veröffentlicht: (2024)
von: Conti, Alessandro, et al.
Veröffentlicht: (2024)
Vocabulary-free Image Classification
von: Conti, Alessandro, et al.
Veröffentlicht: (2023)
von: Conti, Alessandro, et al.
Veröffentlicht: (2023)
Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifiers
von: Guimard, Quentin, et al.
Veröffentlicht: (2025)
von: Guimard, Quentin, et al.
Veröffentlicht: (2025)
Harnessing Large Language Models for Training-free Video Anomaly Detection
von: Zanella, Luca, et al.
Veröffentlicht: (2024)
von: Zanella, Luca, et al.
Veröffentlicht: (2024)
Automatic benchmarking of large multimodal models via iterative experiment programming
von: Conti, Alessandro, et al.
Veröffentlicht: (2024)
von: Conti, Alessandro, et al.
Veröffentlicht: (2024)
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
von: Caldarella, Simone, et al.
Veröffentlicht: (2024)
von: Caldarella, Simone, et al.
Veröffentlicht: (2024)
3D Part Segmentation via Geometric Aggregation of 2D Visual Features
von: Garosi, Marco, et al.
Veröffentlicht: (2024)
von: Garosi, Marco, et al.
Veröffentlicht: (2024)
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering
von: Das, Deepayan, et al.
Veröffentlicht: (2024)
von: Das, Deepayan, et al.
Veröffentlicht: (2024)
Can Text-to-Video Generation help Video-Language Alignment?
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
Training-Free Personalization via Retrieval and Reasoning on Fingerprints
von: Das, Deepayan, et al.
Veröffentlicht: (2025)
von: Das, Deepayan, et al.
Veröffentlicht: (2025)
Training-free Online Video Step Grounding
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
Unlearning Personal Data from a Single Image
von: De Min, Thomas, et al.
Veröffentlicht: (2024)
von: De Min, Thomas, et al.
Veröffentlicht: (2024)
Specificity-aware reinforcement learning for fine-grained open-world classification
von: Angheben, Samuele, et al.
Veröffentlicht: (2026)
von: Angheben, Samuele, et al.
Veröffentlicht: (2026)
Towards Unconstrained Human-Object Interaction
von: Tonini, Francesco, et al.
Veröffentlicht: (2026)
von: Tonini, Francesco, et al.
Veröffentlicht: (2026)
Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection
von: Tonini, Francesco, et al.
Veröffentlicht: (2025)
von: Tonini, Francesco, et al.
Veröffentlicht: (2025)
Test-Time Zero-Shot Temporal Action Localization
von: Liberatori, Benedetta, et al.
Veröffentlicht: (2024)
von: Liberatori, Benedetta, et al.
Veröffentlicht: (2024)
From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition
von: Gentile, Francesco, et al.
Veröffentlicht: (2026)
von: Gentile, Francesco, et al.
Veröffentlicht: (2026)
Test-time Vocabulary Adaptation for Language-driven Object Detection
von: Liu, Mingxuan, et al.
Veröffentlicht: (2025)
von: Liu, Mingxuan, et al.
Veröffentlicht: (2025)
Less is more: Summarizing Patch Tokens for efficient Multi-Label Class-Incremental Learning
von: De Min, Thomas, et al.
Veröffentlicht: (2024)
von: De Min, Thomas, et al.
Veröffentlicht: (2024)
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
von: Huang, Yiran, et al.
Veröffentlicht: (2025)
von: Huang, Yiran, et al.
Veröffentlicht: (2025)
Zero-Shot Temporal Action Localization Through Textual Guidance
von: Liberatori, Benedetta, et al.
Veröffentlicht: (2026)
von: Liberatori, Benedetta, et al.
Veröffentlicht: (2026)
ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
von: Liberatori, Benedetta, et al.
Veröffentlicht: (2025)
von: Liberatori, Benedetta, et al.
Veröffentlicht: (2025)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
von: Guimard, Quentin, et al.
Veröffentlicht: (2026)
von: Guimard, Quentin, et al.
Veröffentlicht: (2026)
Multimodal Large Language Models as Image Classifiers
von: Kisel, Nikita, et al.
Veröffentlicht: (2026)
von: Kisel, Nikita, et al.
Veröffentlicht: (2026)
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
von: Tur, Anil Osman, et al.
Veröffentlicht: (2024)
von: Tur, Anil Osman, et al.
Veröffentlicht: (2024)
Generative Multimodal Models are In-Context Learners
von: Sun, Quan, et al.
Veröffentlicht: (2023)
von: Sun, Quan, et al.
Veröffentlicht: (2023)
GradBias: Unveiling Word Influence on Bias in Text-to-Image Generative Models
von: D'Incà, Moreno, et al.
Veröffentlicht: (2024)
von: D'Incà, Moreno, et al.
Veröffentlicht: (2024)
Kosmos-G: Generating Images in Context with Multimodal Large Language Models
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
Personal Visual Context Learning in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
Linear Model Merging Unlocks Simple and Scalable Multimodal Data Mixture Optimization
von: Berasi, Davide, et al.
Veröffentlicht: (2026)
von: Berasi, Davide, et al.
Veröffentlicht: (2026)
FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
von: Singha, Mainak, et al.
Veröffentlicht: (2025)
von: Singha, Mainak, et al.
Veröffentlicht: (2025)
Democratizing Fine-grained Visual Recognition with Large Language Models
von: Liu, Mingxuan, et al.
Veröffentlicht: (2024)
von: Liu, Mingxuan, et al.
Veröffentlicht: (2024)
CaMML: Context-Aware Multimodal Learner for Large Models
von: Chen, Yixin, et al.
Veröffentlicht: (2024)
von: Chen, Yixin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Compositional Caching for Training-free Open-vocabulary Attribute Detection
von: Garosi, Marco, et al.
Veröffentlicht: (2025) -
On Large Multimodal Models as Open-World Image Classifiers
von: Conti, Alessandro, et al.
Veröffentlicht: (2025) -
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
von: Farina, Matteo, et al.
Veröffentlicht: (2025) -
Frustratingly Easy Test-Time Adaptation of Vision-Language Models
von: Farina, Matteo, et al.
Veröffentlicht: (2024) -
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
von: Berasi, Davide, et al.
Veröffentlicht: (2025)