LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Kaichen, Li, Bo, Zhang, Peiyuan, Pu, Fanyi, Cahyono, Joshua Adrian, Hu, Kairui, Liu, Shuai, Zhang, Yuanhan, Yang, Jingkang, Li, Chunyuan, Liu, Ziwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Otter: A Multi-Modal Model with In-Context Instruction Tuning
von: Li, Bo, et al.
Veröffentlicht: (2023)
von: Li, Bo, et al.
Veröffentlicht: (2023)
WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
Long Context Transfer from Language to Vision
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2024)
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
von: Hu, Kairui, et al.
Veröffentlicht: (2025)
von: Hu, Kairui, et al.
Veröffentlicht: (2025)
LLaVA-OneVision: Easy Visual Task Transfer
von: Li, Bo, et al.
Veröffentlicht: (2024)
von: Li, Bo, et al.
Veröffentlicht: (2024)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
von: Li, Feng, et al.
Veröffentlicht: (2024)
von: Li, Feng, et al.
Veröffentlicht: (2024)
FunQA: Towards Surprising Video Comprehension
von: Xie, Binzhu, et al.
Veröffentlicht: (2023)
von: Xie, Binzhu, et al.
Veröffentlicht: (2023)
Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs
von: Quang, Trung Nguyen, et al.
Veröffentlicht: (2026)
von: Quang, Trung Nguyen, et al.
Veröffentlicht: (2026)
OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
MMSearch-R1: Incentivizing LMMs to Search
von: Wu, Jinming, et al.
Veröffentlicht: (2025)
von: Wu, Jinming, et al.
Veröffentlicht: (2025)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
Can You Trust an LLM with Your Life-Changing Decision? An Investigation into AI High-Stakes Responses
von: Cahyono, Joshua Adrian, et al.
Veröffentlicht: (2025)
von: Cahyono, Joshua Adrian, et al.
Veröffentlicht: (2025)
FileGram: Grounding Agent Personalization in File-System Behavioral Traces
von: Liu, Shuai, et al.
Veröffentlicht: (2026)
von: Liu, Shuai, et al.
Veröffentlicht: (2026)
Octopus: Embodied Vision-Language Programmer from Environmental Feedback
von: Yang, Jingkang, et al.
Veröffentlicht: (2023)
von: Yang, Jingkang, et al.
Veröffentlicht: (2023)
Automated Image Captioning with CNNs and Transformers
von: Cahyono, Joshua Adrian, et al.
Veröffentlicht: (2024)
von: Cahyono, Joshua Adrian, et al.
Veröffentlicht: (2024)
MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding
von: Ye, Zhoutong, et al.
Veröffentlicht: (2025)
von: Ye, Zhoutong, et al.
Veröffentlicht: (2025)
MIBench: Evaluating LMMs on Multimodal Interaction
von: Miao, Yu, et al.
Veröffentlicht: (2026)
von: Miao, Yu, et al.
Veröffentlicht: (2026)
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
von: Zhang, Ruohong, et al.
Veröffentlicht: (2024)
von: Zhang, Ruohong, et al.
Veröffentlicht: (2024)
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
von: Dong, Yuhao, et al.
Veröffentlicht: (2024)
von: Dong, Yuhao, et al.
Veröffentlicht: (2024)
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2025)
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
Generalized Out-of-Distribution Detection: A Survey
von: Yang, Jingkang, et al.
Veröffentlicht: (2021)
von: Yang, Jingkang, et al.
Veröffentlicht: (2021)
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
von: Xu, Zitong, et al.
Veröffentlicht: (2025)
von: Xu, Zitong, et al.
Veröffentlicht: (2025)
EgoLife: Towards Egocentric Life Assistant
von: Yang, Jingkang, et al.
Veröffentlicht: (2025)
von: Yang, Jingkang, et al.
Veröffentlicht: (2025)
Graphic Design with Large Multimodal Model
von: Cheng, Yutao, et al.
Veröffentlicht: (2024)
von: Cheng, Yutao, et al.
Veröffentlicht: (2024)
HippoCamp: Benchmarking Contextual Agents on Personal Computers
von: Yang, Zhe, et al.
Veröffentlicht: (2026)
von: Yang, Zhe, et al.
Veröffentlicht: (2026)
A-Bench: Are LMMs Masters at Evaluating AI-generated Images?
von: Zhang, Zicheng, et al.
Veröffentlicht: (2024)
von: Zhang, Zicheng, et al.
Veröffentlicht: (2024)
OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence
von: Tang, Feilong, et al.
Veröffentlicht: (2026)
von: Tang, Feilong, et al.
Veröffentlicht: (2026)
LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models
von: Li, Haitao, et al.
Veröffentlicht: (2024)
von: Li, Haitao, et al.
Veröffentlicht: (2024)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
von: Li, Can, et al.
Veröffentlicht: (2025)
von: Li, Can, et al.
Veröffentlicht: (2025)
Benchmarking and Analyzing Generative Data for Visual Recognition
von: Li, Bo, et al.
Veröffentlicht: (2023)
von: Li, Bo, et al.
Veröffentlicht: (2023)
AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models
von: Shu, Dong, et al.
Veröffentlicht: (2024)
von: Shu, Dong, et al.
Veröffentlicht: (2024)
AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs
von: Xia, Shuhan, et al.
Veröffentlicht: (2025)
von: Xia, Shuhan, et al.
Veröffentlicht: (2025)
Taxonomy-based CheckList for Large Language Model Evaluation
von: Zhang, Damin
Veröffentlicht: (2023)
von: Zhang, Damin
Veröffentlicht: (2023)
MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Otter: A Multi-Modal Model with In-Context Instruction Tuning
von: Li, Bo, et al.
Veröffentlicht: (2023) -
WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024) -
Long Context Transfer from Language to Vision
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2024) -
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
von: Hu, Kairui, et al.
Veröffentlicht: (2025) -
LLaVA-OneVision: Easy Visual Task Transfer
von: Li, Bo, et al.
Veröffentlicht: (2024)