Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Xinyu, Chen, Zhijian, Lan, Kai, Ma, Lixin, Ding, Shengyuan, Liang, Yingji, Zhao, Xiangyu, Wen, Farong, Zhang, Zicheng, Zhang, Guofeng, Duan, Haodong, Chen, Kai, Lin, Dahua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
MMBench: Is Your Multi-modal Model an All-around Player?
by: Liu, Yuan, et al.
Published: (2023)
by: Liu, Yuan, et al.
Published: (2023)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
by: Zhuo, Jingming, et al.
Published: (2024)
by: Zhuo, Jingming, et al.
Published: (2024)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024)
by: Qiao, Yuxuan, et al.
Published: (2024)
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
by: Li, Xiaozhe, et al.
Published: (2025)
by: Li, Xiaozhe, et al.
Published: (2025)
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
by: Wang, Chonghua, et al.
Published: (2024)
by: Wang, Chonghua, et al.
Published: (2024)
Information Density Principle for MLLM Benchmarks
by: Li, Chunyi, et al.
Published: (2025)
by: Li, Chunyi, et al.
Published: (2025)
NP-Engine: Empowering Optimization Reasoning in Large Language Models with Verifiable Synthetic NP Problems
by: Li, Xiaozhe, et al.
Published: (2025)
by: Li, Xiaozhe, et al.
Published: (2025)
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
by: Shen, Ye, et al.
Published: (2025)
by: Shen, Ye, et al.
Published: (2025)
Redundancy Principles for MLLMs Benchmarks
by: Zhang, Zicheng, et al.
Published: (2025)
by: Zhang, Zicheng, et al.
Published: (2025)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
by: Wang, Xuehui, et al.
Published: (2025)
by: Wang, Xuehui, et al.
Published: (2025)
Improve MLLM Benchmark Efficiency through Interview
by: Wen, Farong, et al.
Published: (2025)
by: Wen, Farong, et al.
Published: (2025)
MM-IFEngine: Towards Multimodal Instruction Following
by: Ding, Shengyuan, et al.
Published: (2025)
by: Ding, Shengyuan, et al.
Published: (2025)
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
by: Zang, Yuhang, et al.
Published: (2025)
by: Zang, Yuhang, et al.
Published: (2025)
SPARK: Synergistic Policy And Reward Co-Evolving Framework
by: Liu, Ziyu, et al.
Published: (2025)
by: Liu, Ziyu, et al.
Published: (2025)
GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI
by: Chen, Pengcheng, et al.
Published: (2024)
by: Chen, Pengcheng, et al.
Published: (2024)
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
by: Ding, Shengyuan, et al.
Published: (2025)
by: Ding, Shengyuan, et al.
Published: (2025)
Metaphor Creation: A Measure of Creativity or Intelligence?
by: Débora Pereira de Barros
Published: (2010)
by: Débora Pereira de Barros
Published: (2010)
GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing
by: Xiao, Aoran, et al.
Published: (2026)
by: Xiao, Aoran, et al.
Published: (2026)
Affordance Benchmark for MLLMs
by: Wang, Junying, et al.
Published: (2025)
by: Wang, Junying, et al.
Published: (2025)
Visual-ERM: Reward Modeling for Visual Equivalence
by: Liu, Ziyu, et al.
Published: (2026)
by: Liu, Ziyu, et al.
Published: (2026)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
by: Ding, Shuangrui, et al.
Published: (2026)
by: Ding, Shuangrui, et al.
Published: (2026)
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
by: Wu, Diankun, et al.
Published: (2025)
by: Wu, Diankun, et al.
Published: (2025)
Formal Semantic Control over Language Models
by: Zhang, Yingji
Published: (2026)
by: Zhang, Yingji
Published: (2026)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
by: Wu, Yuhang, et al.
Published: (2024)
by: Wu, Yuhang, et al.
Published: (2024)
From Perceived School Climate to Creativity Performance: The Serial Multiple Mediation of Creative Self‐Efficacy and Creativity Motivation
by: Wu‐jing He, et al.
Published: (2025)
by: Wu‐jing He, et al.
Published: (2025)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
by: Liu, Hongwei, et al.
Published: (2024)
by: Liu, Hongwei, et al.
Published: (2024)
AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process
by: Zhang, Xintong, et al.
Published: (2026)
by: Zhang, Xintong, et al.
Published: (2026)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
by: Li, Mo, et al.
Published: (2024)
by: Li, Mo, et al.
Published: (2024)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
by: Li, Jidong, et al.
Published: (2025)
by: Li, Jidong, et al.
Published: (2025)
The Ever-Evolving Science Exam
by: Wang, Junying, et al.
Published: (2025)
by: Wang, Junying, et al.
Published: (2025)
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
by: Wang, Junying, et al.
Published: (2025)
by: Wang, Junying, et al.
Published: (2025)
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
by: Cao, Maosong, et al.
Published: (2024)
by: Cao, Maosong, et al.
Published: (2024)
Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence
by: Zheng, Yufei, et al.
Published: (2026)
by: Zheng, Yufei, et al.
Published: (2026)
CriticEval: Evaluating Large Language Model as Critic
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
Creativity, Innovation, and Knowledge Creation in the Era of Generative Artificial Intelligence: Challenging Human Intelligence
by: Namita Jain, et al.
Published: (2025)
by: Namita Jain, et al.
Published: (2025)
Similar Items
-
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024) -
MMBench: Is Your Multi-modal Model an All-around Player?
by: Liu, Yuan, et al.
Published: (2023) -
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
by: Zhuo, Jingming, et al.
Published: (2024) -
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024) -
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs
by: Li, Xiaozhe, et al.
Published: (2026)