MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Weihao, Yang, Zhengyuan, Ren, Lingfeng, Li, Linjie, Wang, Jianfeng, Lin, Kevin, Lin, Chung-Ching, Liu, Zicheng, Wang, Lijuan, Wang, Xinchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
von: Yang, Zhengyuan, et al.
Veröffentlicht: (2023)
von: Yang, Zhengyuan, et al.
Veröffentlicht: (2023)
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
von: Zhai, Yuanhao, et al.
Veröffentlicht: (2024)
von: Zhai, Yuanhao, et al.
Veröffentlicht: (2024)
DisCo: Disentangled Control for Realistic Human Dance Generation
von: Wang, Tan, et al.
Veröffentlicht: (2023)
von: Wang, Tan, et al.
Veröffentlicht: (2023)
Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
von: Zhai, Yuanhao, et al.
Veröffentlicht: (2024)
von: Zhai, Yuanhao, et al.
Veröffentlicht: (2024)
NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
von: Ren, Lingfeng, et al.
Veröffentlicht: (2026)
von: Ren, Lingfeng, et al.
Veröffentlicht: (2026)
Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2025)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2025)
Bring Metric Functions into Diffusion Models
von: An, Jie, et al.
Veröffentlicht: (2024)
von: An, Jie, et al.
Veröffentlicht: (2024)
GenXD: Generating Any 3D and 4D Scenes
von: Zhao, Yuyang, et al.
Veröffentlicht: (2024)
von: Zhao, Yuyang, et al.
Veröffentlicht: (2024)
V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
von: Zheng, Xiangxi, et al.
Veröffentlicht: (2025)
von: Zheng, Xiangxi, et al.
Veröffentlicht: (2025)
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
LiVOS: Light Video Object Segmentation with Gated Linear Matching
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
von: Yan, An, et al.
Veröffentlicht: (2024)
von: Yan, An, et al.
Veröffentlicht: (2024)
SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
von: Wang, Xiyao, et al.
Veröffentlicht: (2025)
von: Wang, Xiyao, et al.
Veröffentlicht: (2025)
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
Measurement of LLM's Philosophies of Human Nature
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
von: Hong, Yining, et al.
Veröffentlicht: (2024)
von: Hong, Yining, et al.
Veröffentlicht: (2024)
Audio-Aware Large Language Models as Judges for Speaking Styles
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition
von: Qiu, Jielin, et al.
Veröffentlicht: (2024)
von: Qiu, Jielin, et al.
Veröffentlicht: (2024)
Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2025)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2025)
STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
Tuning Timestep-Distilled Diffusion Model Using Pairwise Sample Optimization
von: Miao, Zichen, et al.
Veröffentlicht: (2024)
von: Miao, Zichen, et al.
Veröffentlicht: (2024)
EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2024)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2024)
Computer-Use Agents as Judges for Generative User Interface
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
von: He, Xuehai, et al.
Veröffentlicht: (2024)
von: He, Xuehai, et al.
Veröffentlicht: (2024)
Attention Prompting on Image for Large Vision-Language Models
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
von: li, Bonan, et al.
Veröffentlicht: (2025)
von: li, Bonan, et al.
Veröffentlicht: (2025)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
MambaOut: Do We Really Need Mamba for Vision?
von: Yu, Weihao, et al.
Veröffentlicht: (2024)
von: Yu, Weihao, et al.
Veröffentlicht: (2024)
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
von: Wang, Xiyao, et al.
Veröffentlicht: (2025)
von: Wang, Xiyao, et al.
Veröffentlicht: (2025)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
von: Yu, Weihao, et al.
Veröffentlicht: (2023) -
Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
von: Yang, Zhengyuan, et al.
Veröffentlicht: (2023) -
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
von: Ni, Minheng, et al.
Veröffentlicht: (2025) -
IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
von: Zhai, Yuanhao, et al.
Veröffentlicht: (2024) -
DisCo: Disentangled Control for Realistic Human Dance Generation
von: Wang, Tan, et al.
Veröffentlicht: (2023)