Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography
Fuente:
arXiv
Salvato in:
| Autori principali: | Spinaci, Gianmarco, Klic, Lukas, Colavizza, Giovanni |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Benchmarking Large Language Models for Handwritten Text Recognition
di: Crosilla, Giorgia, et al.
Pubblicazione: (2025)
di: Crosilla, Giorgia, et al.
Pubblicazione: (2025)
FADE: Few-shot/zero-shot Anomaly Detection Engine using Large Vision-Language Model
di: Li, Yuanwei, et al.
Pubblicazione: (2024)
di: Li, Yuanwei, et al.
Pubblicazione: (2024)
Making Large Vision Language Models to be Good Few-shot Learners
di: Liu, Fan, et al.
Pubblicazione: (2024)
di: Liu, Fan, et al.
Pubblicazione: (2024)
Few-shot Adaptation of Medical Vision-Language Models
di: Shakeri, Fereshteh, et al.
Pubblicazione: (2024)
di: Shakeri, Fereshteh, et al.
Pubblicazione: (2024)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
di: Aklilu, Josiah, et al.
Pubblicazione: (2024)
di: Aklilu, Josiah, et al.
Pubblicazione: (2024)
Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners
di: He, Xuehai, et al.
Pubblicazione: (2023)
di: He, Xuehai, et al.
Pubblicazione: (2023)
What Makes Good Few-shot Examples for Vision-Language Models?
di: Guo, Zhaojun, et al.
Pubblicazione: (2024)
di: Guo, Zhaojun, et al.
Pubblicazione: (2024)
Meta-Adapter: An Online Few-shot Learner for Vision-Language Model
di: Cheng, Cheng, et al.
Pubblicazione: (2023)
di: Cheng, Cheng, et al.
Pubblicazione: (2023)
DiffCLIP: Few-shot Language-driven Multimodal Classifier
di: Zhang, Jiaqing, et al.
Pubblicazione: (2024)
di: Zhang, Jiaqing, et al.
Pubblicazione: (2024)
Vision and Language Reference Prompt into SAM for Few-shot Segmentation
di: Sakurai, Kosuke, et al.
Pubblicazione: (2025)
di: Sakurai, Kosuke, et al.
Pubblicazione: (2025)
Selective Vision-Language Subspace Projection for Few-shot CLIP
di: Zhu, Xingyu, et al.
Pubblicazione: (2024)
di: Zhu, Xingyu, et al.
Pubblicazione: (2024)
A Chain-of-Thought Subspace Meta-Learning for Few-shot Image Captioning with Large Vision and Language Models
di: Huang, Hao, et al.
Pubblicazione: (2025)
di: Huang, Hao, et al.
Pubblicazione: (2025)
Envisioning Class Entity Reasoning by Large Language Models for Few-shot Learning
di: Liu, Mushui, et al.
Pubblicazione: (2024)
di: Liu, Mushui, et al.
Pubblicazione: (2024)
Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-shot Semantic Segmentation
di: Liu, Jie, et al.
Pubblicazione: (2025)
di: Liu, Jie, et al.
Pubblicazione: (2025)
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
di: An, Zhaochong, et al.
Pubblicazione: (2025)
di: An, Zhaochong, et al.
Pubblicazione: (2025)
Label Propagation for Zero-shot Classification with Vision-Language Models
di: Stojnić, Vladan, et al.
Pubblicazione: (2024)
di: Stojnić, Vladan, et al.
Pubblicazione: (2024)
Boosting Audio-visual Zero-shot Learning with Large Language Models
di: Chen, Haoxing, et al.
Pubblicazione: (2023)
di: Chen, Haoxing, et al.
Pubblicazione: (2023)
Manifold Induced Biases for Zero-shot and Few-shot Detection of Generated Images
di: Brokman, Jonathan, et al.
Pubblicazione: (2025)
di: Brokman, Jonathan, et al.
Pubblicazione: (2025)
Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection
di: Tang, Lv, et al.
Pubblicazione: (2023)
di: Tang, Lv, et al.
Pubblicazione: (2023)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
di: Xu, Yifang, et al.
Pubblicazione: (2025)
di: Xu, Yifang, et al.
Pubblicazione: (2025)
Zero-shot Generalizable Incremental Learning for Vision-Language Object Detection
di: Deng, Jieren, et al.
Pubblicazione: (2024)
di: Deng, Jieren, et al.
Pubblicazione: (2024)
Zero-shot image privacy classification with Vision-Language Models
di: Baia, Alina Elena, et al.
Pubblicazione: (2025)
di: Baia, Alina Elena, et al.
Pubblicazione: (2025)
Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
di: Ding, Guodong, et al.
Pubblicazione: (2026)
di: Ding, Guodong, et al.
Pubblicazione: (2026)
FLIER: Few-shot Language Image Models Embedded with Latent Representations
di: Zhou, Zhinuo, et al.
Pubblicazione: (2024)
di: Zhou, Zhinuo, et al.
Pubblicazione: (2024)
Zero-shot Vision-Language Reranking for Cross-View Geolocalization
di: Erzurumlu, Yunus Talha, et al.
Pubblicazione: (2026)
di: Erzurumlu, Yunus Talha, et al.
Pubblicazione: (2026)
Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
di: Bui, Phuoc-Nguyen, et al.
Pubblicazione: (2025)
di: Bui, Phuoc-Nguyen, et al.
Pubblicazione: (2025)
Zero-shot Hazard Identification in Autonomous Driving: A Case Study on the COOOL Benchmark
di: Picek, Lukas, et al.
Pubblicazione: (2024)
di: Picek, Lukas, et al.
Pubblicazione: (2024)
LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?
di: Li, Bangyan, et al.
Pubblicazione: (2025)
di: Li, Bangyan, et al.
Pubblicazione: (2025)
Unified Language-driven Zero-shot Domain Adaptation
di: Yang, Senqiao, et al.
Pubblicazione: (2024)
di: Yang, Senqiao, et al.
Pubblicazione: (2024)
Towards Zero-shot Human-Object Interaction Detection via Vision-Language Integration
di: Xue, Weiying, et al.
Pubblicazione: (2024)
di: Xue, Weiying, et al.
Pubblicazione: (2024)
RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models
di: Li, Junjie, et al.
Pubblicazione: (2025)
di: Li, Junjie, et al.
Pubblicazione: (2025)
Unbiased Semantic Decoding with Vision Foundation Models for Few-shot Segmentation
di: Wang, Jin, et al.
Pubblicazione: (2025)
di: Wang, Jin, et al.
Pubblicazione: (2025)
Vocabulary-free few-shot learning for Vision-Language Models
di: Zanella, Maxime, et al.
Pubblicazione: (2025)
di: Zanella, Maxime, et al.
Pubblicazione: (2025)
VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
di: Wu, Pengying, et al.
Pubblicazione: (2024)
di: Wu, Pengying, et al.
Pubblicazione: (2024)
Zero-shot detection of buildings in mobile LiDAR using Language Vision Model
di: Goo, June Moh, et al.
Pubblicazione: (2024)
di: Goo, June Moh, et al.
Pubblicazione: (2024)
Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain Adaptation
di: Liu, Kuanghong, et al.
Pubblicazione: (2025)
di: Liu, Kuanghong, et al.
Pubblicazione: (2025)
Language-Inspired Relation Transfer for Few-shot Class-Incremental Learning
di: Zhao, Yifan, et al.
Pubblicazione: (2025)
di: Zhao, Yifan, et al.
Pubblicazione: (2025)
MAKE: Multi-Aspect Knowledge-Enhanced Vision-Language Pretraining for Zero-shot Dermatological Assessment
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
di: Tur, Anil Osman, et al.
Pubblicazione: (2024)
di: Tur, Anil Osman, et al.
Pubblicazione: (2024)
Few-shot Object Localization
di: Ren, Yunhan, et al.
Pubblicazione: (2024)
di: Ren, Yunhan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Benchmarking Large Language Models for Handwritten Text Recognition
di: Crosilla, Giorgia, et al.
Pubblicazione: (2025) -
FADE: Few-shot/zero-shot Anomaly Detection Engine using Large Vision-Language Model
di: Li, Yuanwei, et al.
Pubblicazione: (2024) -
Making Large Vision Language Models to be Good Few-shot Learners
di: Liu, Fan, et al.
Pubblicazione: (2024) -
Few-shot Adaptation of Medical Vision-Language Models
di: Shakeri, Fereshteh, et al.
Pubblicazione: (2024) -
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
di: Aklilu, Josiah, et al.
Pubblicazione: (2024)