Revisiting Multi-Modal LLM Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Jian, Srivastava, Shikhar, Chen, Junyu, Shrestha, Robik, Acharya, Manoj, Kafle, Kushal, Kanan, Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
Are Bias Mitigation Techniques for Deep Learning Effective?
by: Shrestha, Robik, et al.
Published: (2021)
by: Shrestha, Robik, et al.
Published: (2021)
Improving Multimodal Large Language Models Using Continual Learning
by: Srivastava, Shikhar, et al.
Published: (2024)
by: Srivastava, Shikhar, et al.
Published: (2024)
BloomVQA: Assessing Hierarchical Multi-modal Comprehension
by: Gong, Yunye, et al.
Published: (2023)
by: Gong, Yunye, et al.
Published: (2023)
Learning to Evaluate the Artness of AI-generated Images
by: Chen, Junyu, et al.
Published: (2023)
by: Chen, Junyu, et al.
Published: (2023)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
by: Slyman, Eric, et al.
Published: (2025)
by: Slyman, Eric, et al.
Published: (2025)
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
by: Padhi, Trilok, et al.
Published: (2025)
by: Padhi, Trilok, et al.
Published: (2025)
GRASP: A Rehearsal Policy for Efficient Online Continual Learning
by: Harun, Md Yousuf, et al.
Published: (2023)
by: Harun, Md Yousuf, et al.
Published: (2023)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025)
by: Zang, Yuan, et al.
Published: (2025)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
Unlocking ImageNet's Multi-Object Nature: Automated Large-Scale Multilabel Annotation
by: Chen, Junyu, et al.
Published: (2026)
by: Chen, Junyu, et al.
Published: (2026)
INSIGHT: Explainable Weakly-Supervised Medical Image Analysis
by: Zhang, Wenbo, et al.
Published: (2024)
by: Zhang, Wenbo, et al.
Published: (2024)
Multi-Modal Language Models as Text-to-Image Model Evaluators
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
OccamNets: Mitigating Dataset Bias by Favoring Simpler Hypotheses
by: Shrestha, Robik, et al.
Published: (2022)
by: Shrestha, Robik, et al.
Published: (2022)
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
by: Slyman, Eric, et al.
Published: (2024)
by: Slyman, Eric, et al.
Published: (2024)
From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
by: Ye, Hanrong, et al.
Published: (2025)
by: Ye, Hanrong, et al.
Published: (2025)
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
by: Qin, Libo, et al.
Published: (2024)
by: Qin, Libo, et al.
Published: (2024)
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
by: Zhao, Haozhe, et al.
Published: (2023)
by: Zhao, Haozhe, et al.
Published: (2023)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
by: Yao, Huanjin, et al.
Published: (2025)
by: Yao, Huanjin, et al.
Published: (2025)
SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
by: Chen, Yi-Chia, et al.
Published: (2024)
by: Chen, Yi-Chia, et al.
Published: (2024)
Core Knowledge Deficits in Multi-Modal Language Models
by: Li, Yijiang, et al.
Published: (2024)
by: Li, Yijiang, et al.
Published: (2024)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
APLe: Token-Wise Adaptive for Multi-Modal Prompt Learning
by: Cao, Guiming, et al.
Published: (2024)
by: Cao, Guiming, et al.
Published: (2024)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
by: Kou, Siqi, et al.
Published: (2024)
by: Kou, Siqi, et al.
Published: (2024)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
by: Zou, Chengke, et al.
Published: (2024)
by: Zou, Chengke, et al.
Published: (2024)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
by: Li, Zejun, et al.
Published: (2024)
by: Li, Zejun, et al.
Published: (2024)
LLM Augmented LLMs: Expanding Capabilities through Composition
by: Bansal, Rachit, et al.
Published: (2024)
by: Bansal, Rachit, et al.
Published: (2024)
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
by: Zhong, Yiwu, et al.
Published: (2024)
by: Zhong, Yiwu, et al.
Published: (2024)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
Revisiting the Role of Language Priors in Vision-Language Models
by: Lin, Zhiqiu, et al.
Published: (2023)
by: Lin, Zhiqiu, et al.
Published: (2023)
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
ML-Mamba: Efficient Multi-Modal Large Language Model Utilizing Mamba-2
by: Huang, Wenjun, et al.
Published: (2024)
by: Huang, Wenjun, et al.
Published: (2024)
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
by: Oh, Youngtaek, et al.
Published: (2024)
by: Oh, Youngtaek, et al.
Published: (2024)
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
by: Li, Dingming, et al.
Published: (2025)
by: Li, Dingming, et al.
Published: (2025)
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
How Well Do Multi-modal LLMs Interpret CT Scans? An Auto-Evaluation Framework for Analyses
by: Zhu, Qingqing, et al.
Published: (2024)
by: Zhu, Qingqing, et al.
Published: (2024)
Region-R1: Reinforcing Query-Side Region Cropping for Multi-Modal Re-Ranking
by: Hu, Chan-Wei, et al.
Published: (2026)
by: Hu, Chan-Wei, et al.
Published: (2026)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
by: Zhang, Wanpeng, et al.
Published: (2024)
by: Zhang, Wanpeng, et al.
Published: (2024)
Similar Items
-
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020) -
Are Bias Mitigation Techniques for Deep Learning Effective?
by: Shrestha, Robik, et al.
Published: (2021) -
Improving Multimodal Large Language Models Using Continual Learning
by: Srivastava, Shikhar, et al.
Published: (2024) -
BloomVQA: Assessing Hierarchical Multi-modal Comprehension
by: Gong, Yunye, et al.
Published: (2023) -
Learning to Evaluate the Artness of AI-generated Images
by: Chen, Junyu, et al.
Published: (2023)