MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Wulin, Zhang, Yi-Fan, Fu, Chaoyou, Shi, Yang, Nie, Bingyan, Chen, Hongkai, Zhang, Zhang, Wang, Liang, Tan, Tieniu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2024)
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023)
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
von: Fu, Chaoyou, et al.
Veröffentlicht: (2026)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2026)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
von: Shi, Yang, et al.
Veröffentlicht: (2025)
von: Shi, Yang, et al.
Veröffentlicht: (2025)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
von: Gan, Ziliang, et al.
Veröffentlicht: (2024)
von: Gan, Ziliang, et al.
Veröffentlicht: (2024)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
Incomplete Multi-view Multi-label Classification via a Dual-level Contrastive Learning Framework
von: Nie, Bingyan, et al.
Veröffentlicht: (2024)
von: Nie, Bingyan, et al.
Veröffentlicht: (2024)
Human Image Generation: A Comprehensive Survey
von: Jia, Zhen, et al.
Veröffentlicht: (2022)
von: Jia, Zhen, et al.
Veröffentlicht: (2022)
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
von: Shi, Mengqi, et al.
Veröffentlicht: (2026)
von: Shi, Mengqi, et al.
Veröffentlicht: (2026)
Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
von: Liu, Yuansen, et al.
Veröffentlicht: (2025)
von: Liu, Yuansen, et al.
Veröffentlicht: (2025)
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
von: Li, Lijiang, et al.
Veröffentlicht: (2026)
von: Li, Lijiang, et al.
Veröffentlicht: (2026)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
Multi-View Factorizing and Disentangling: A Novel Framework for Incomplete Multi-View Multi-Label Classification
von: Xie, Wulin, et al.
Veröffentlicht: (2025)
von: Xie, Wulin, et al.
Veröffentlicht: (2025)
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2024)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
von: Zhao, Shanshan, et al.
Veröffentlicht: (2025)
von: Zhao, Shanshan, et al.
Veröffentlicht: (2025)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
von: Li, Yi, et al.
Veröffentlicht: (2025)
von: Li, Yi, et al.
Veröffentlicht: (2025)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
von: Zhang, Fan, et al.
Veröffentlicht: (2025)
von: Zhang, Fan, et al.
Veröffentlicht: (2025)
PersonaVLM: Long-Term Personalized Multimodal LLMs
von: Nie, Chang, et al.
Veröffentlicht: (2026)
von: Nie, Chang, et al.
Veröffentlicht: (2026)
UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
von: Zhu, Yijie, et al.
Veröffentlicht: (2025)
von: Zhu, Yijie, et al.
Veröffentlicht: (2025)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
von: Liu, Zeyu, et al.
Veröffentlicht: (2026)
von: Liu, Zeyu, et al.
Veröffentlicht: (2026)
DreamWorld: Unified World Modeling in Video Generation
von: Tan, Boming, et al.
Veröffentlicht: (2026)
von: Tan, Boming, et al.
Veröffentlicht: (2026)
CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2026)
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2026)
A Unified Image-Dense Annotation Generation Model for Underwater Scenes
von: Lin, Hongkai, et al.
Veröffentlicht: (2025)
von: Lin, Hongkai, et al.
Veröffentlicht: (2025)
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
von: Tian, Changyao, et al.
Veröffentlicht: (2026)
von: Tian, Changyao, et al.
Veröffentlicht: (2026)
Unified Reward Model for Multimodal Understanding and Generation
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
von: Zhang, Huichao, et al.
Veröffentlicht: (2026)
von: Zhang, Huichao, et al.
Veröffentlicht: (2026)
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
von: Chen, Zhihong, et al.
Veröffentlicht: (2025)
von: Chen, Zhihong, et al.
Veröffentlicht: (2025)
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
UEval: A Benchmark for Unified Multimodal Generation
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Aligning Multimodal LLM with Human Preference: A Survey
von: Yu, Tao, et al.
Veröffentlicht: (2025)
von: Yu, Tao, et al.
Veröffentlicht: (2025)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
von: Wu, Size, et al.
Veröffentlicht: (2025)
von: Wu, Size, et al.
Veröffentlicht: (2025)
Planning with Unified Multimodal Models
von: Sun, Yihao, et al.
Veröffentlicht: (2025)
von: Sun, Yihao, et al.
Veröffentlicht: (2025)
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2024) -
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023) -
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
von: Fu, Chaoyou, et al.
Veröffentlicht: (2026) -
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
von: Shi, Yang, et al.
Veröffentlicht: (2025) -
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)