FewMMBench: A Benchmark for Multimodal Few-Shot Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Dogan, Mustafa, Kesen, Ilker, Calixto, Iacer, Erdem, Aykut, Erdem, Erkut |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
di: Dogan, Mustafa, et al.
Pubblicazione: (2024)
di: Dogan, Mustafa, et al.
Pubblicazione: (2024)
DeVisE: Behavioral Testing of Medical Large Language Models
di: Tagliabue, Camila Zurdo, et al.
Pubblicazione: (2025)
di: Tagliabue, Camila Zurdo, et al.
Pubblicazione: (2025)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2024)
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2024)
Sequential Compositional Generalization in Multimodal Models
di: Yagcioglu, Semih, et al.
Pubblicazione: (2024)
di: Yagcioglu, Semih, et al.
Pubblicazione: (2024)
Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding
di: Vural, Hatice Merve, et al.
Pubblicazione: (2026)
di: Vural, Hatice Merve, et al.
Pubblicazione: (2026)
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
di: Ercan, Burak, et al.
Pubblicazione: (2023)
di: Ercan, Burak, et al.
Pubblicazione: (2023)
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
di: Sanli, Enes, et al.
Pubblicazione: (2025)
di: Sanli, Enes, et al.
Pubblicazione: (2025)
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
di: Karanfil, Enes, et al.
Pubblicazione: (2025)
di: Karanfil, Enes, et al.
Pubblicazione: (2025)
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
di: Ercan, Burak, et al.
Pubblicazione: (2024)
di: Ercan, Burak, et al.
Pubblicazione: (2024)
Mind the Gap: Benchmarking LLM Uncertainty and Calibration with Specialty-Aware Clinical QA and Reasoning-Based Behavioural Features
di: Testoni, Alberto, et al.
Pubblicazione: (2025)
di: Testoni, Alberto, et al.
Pubblicazione: (2025)
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
di: Bond, Andrew, et al.
Pubblicazione: (2026)
di: Bond, Andrew, et al.
Pubblicazione: (2026)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
di: Ercan, Burak, et al.
Pubblicazione: (2023)
di: Ercan, Burak, et al.
Pubblicazione: (2023)
Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA
di: Testoni, Alberto, et al.
Pubblicazione: (2026)
di: Testoni, Alberto, et al.
Pubblicazione: (2026)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
di: Bond, Andrew, et al.
Pubblicazione: (2025)
di: Bond, Andrew, et al.
Pubblicazione: (2025)
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
di: Ekin, Yigit, et al.
Pubblicazione: (2024)
di: Ekin, Yigit, et al.
Pubblicazione: (2024)
AnyMatch -- Efficient Zero-Shot Entity Matching with a Small Language Model
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking
di: Mishra, Nishant, et al.
Pubblicazione: (2025)
di: Mishra, Nishant, et al.
Pubblicazione: (2025)
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
di: Çapuk, Hakan, et al.
Pubblicazione: (2025)
di: Çapuk, Hakan, et al.
Pubblicazione: (2025)
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
di: Kizil, Muhammed Burak, et al.
Pubblicazione: (2025)
di: Kizil, Muhammed Burak, et al.
Pubblicazione: (2025)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
di: Cokelek, Mert, et al.
Pubblicazione: (2025)
di: Cokelek, Mert, et al.
Pubblicazione: (2025)
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
di: Anees, Abdul Basit, et al.
Pubblicazione: (2024)
di: Anees, Abdul Basit, et al.
Pubblicazione: (2024)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
di: Ali, Moayed Haji, et al.
Pubblicazione: (2023)
di: Ali, Moayed Haji, et al.
Pubblicazione: (2023)
SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
di: Biner, Burak Can, et al.
Pubblicazione: (2024)
di: Biner, Burak Can, et al.
Pubblicazione: (2024)
Object and Relation Centric Representations for Push Effect Prediction
di: Tekden, Ahmet E., et al.
Pubblicazione: (2021)
di: Tekden, Ahmet E., et al.
Pubblicazione: (2021)
Benchmarking Open-Source Large Language Models for Persian in Zero-Shot and Few-Shot Learning
di: Cherakhloo, Mahdi, et al.
Pubblicazione: (2025)
di: Cherakhloo, Mahdi, et al.
Pubblicazione: (2025)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
di: Kizil, Muhammed Burak, et al.
Pubblicazione: (2026)
di: Kizil, Muhammed Burak, et al.
Pubblicazione: (2026)
What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs
di: Yan, Xinlan, et al.
Pubblicazione: (2025)
di: Yan, Xinlan, et al.
Pubblicazione: (2025)
A Few Hypocrites: Few-Shot Learning and Subtype Definitions for Detecting Hypocrisy Accusations in Online Climate Change Debates
di: Corral, Paulina Garcia, et al.
Pubblicazione: (2024)
di: Corral, Paulina Garcia, et al.
Pubblicazione: (2024)
Active Few-Shot Learning for Text Classification
di: Ahmadnia, Saeed, et al.
Pubblicazione: (2025)
di: Ahmadnia, Saeed, et al.
Pubblicazione: (2025)
Evaluation of Few-Shot Learning for Classification Tasks in the Polish Language
di: Hadeliya, Tsimur, et al.
Pubblicazione: (2024)
di: Hadeliya, Tsimur, et al.
Pubblicazione: (2024)
In-Context Learning Distillation for Efficient Few-Shot Fine-Tuning
di: Duan, Yifei, et al.
Pubblicazione: (2024)
di: Duan, Yifei, et al.
Pubblicazione: (2024)
Retrieving Versus Understanding Extractive Evidence in Few-Shot Learning
di: Elbakian, Karl, et al.
Pubblicazione: (2025)
di: Elbakian, Karl, et al.
Pubblicazione: (2025)
In-Context Learning for Few-Shot Nested Named Entity Recognition
di: Zhang, Meishan, et al.
Pubblicazione: (2024)
di: Zhang, Meishan, et al.
Pubblicazione: (2024)
Dual Modality-Aware Gated Prompt Tuning for Few-Shot Multimodal Sarcasm Detection
di: Jana, Soumyadeep, et al.
Pubblicazione: (2025)
di: Jana, Soumyadeep, et al.
Pubblicazione: (2025)
Class-Incremental Few-Shot Event Detection
di: Zhao, Kailin, et al.
Pubblicazione: (2024)
di: Zhao, Kailin, et al.
Pubblicazione: (2024)
Language Models are Few-Shot Graders
di: Zhao, Chenyan, et al.
Pubblicazione: (2025)
di: Zhao, Chenyan, et al.
Pubblicazione: (2025)
Prompt Tuning for Few-Shot Continual Learning Named Entity Recognition
di: Ren, Zhe
Pubblicazione: (2025)
di: Ren, Zhe
Pubblicazione: (2025)
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2021)
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2021)
Cross-Modal Augmentation for Few-Shot Multimodal Fake News Detection
di: Jiang, Ye, et al.
Pubblicazione: (2024)
di: Jiang, Ye, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
di: Dogan, Mustafa, et al.
Pubblicazione: (2024) -
DeVisE: Behavioral Testing of Medical Large Language Models
di: Tagliabue, Camila Zurdo, et al.
Pubblicazione: (2025) -
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025) -
Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2024) -
Sequential Compositional Generalization in Multimodal Models
di: Yagcioglu, Semih, et al.
Pubblicazione: (2024)