OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Pengfei, Peng, Xiaopeng, Song, Jiajun, Li, Chuanhao, Xu, Zhaopan, Yang, Yue, Guo, Ziyao, Zhang, Hao, Lin, Yuqi, He, Yefei, Zhao, Lirui, Liu, Shuo, Li, Tianhua, Xie, Yuxuan, Chang, Xiaojun, Qiao, Yu, Shao, Wenqi, Zhang, Kaipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
von: Xu, Zhaopan, et al.
Veröffentlicht: (2025)
von: Xu, Zhaopan, et al.
Veröffentlicht: (2025)
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
von: Ai, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ai, Jiaxin, et al.
Veröffentlicht: (2025)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
von: Li, Chuanhao, et al.
Veröffentlicht: (2024)
von: Li, Chuanhao, et al.
Veröffentlicht: (2024)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
von: Feng, Yukang, et al.
Veröffentlicht: (2025)
von: Feng, Yukang, et al.
Veröffentlicht: (2025)
DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
von: Xu, Zhaopan, et al.
Veröffentlicht: (2025)
von: Xu, Zhaopan, et al.
Veröffentlicht: (2025)
ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification
von: He, Yefei, et al.
Veröffentlicht: (2024)
von: He, Yefei, et al.
Veröffentlicht: (2024)
AI Idea Bench 2025: AI Research Idea Generation Benchmark
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
AgentCPM-Report: Interleaving Drafting and Deepening for Open-Ended Deep Research
von: Li, Yishan, et al.
Veröffentlicht: (2026)
von: Li, Yishan, et al.
Veröffentlicht: (2026)
ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks
von: Yang, Yan, et al.
Veröffentlicht: (2025)
von: Yang, Yan, et al.
Veröffentlicht: (2025)
Scrambling Enabled Entropy Accumulation in Open Quantum Systems
von: Zhang, Yuke, et al.
Veröffentlicht: (2025)
von: Zhang, Yuke, et al.
Veröffentlicht: (2025)
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
von: Chern, Ethan, et al.
Veröffentlicht: (2024)
von: Chern, Ethan, et al.
Veröffentlicht: (2024)
Improving Autoregressive Image Generation through Coarse-to-Fine Token Prediction
von: Guo, Ziyao, et al.
Veröffentlicht: (2025)
von: Guo, Ziyao, et al.
Veröffentlicht: (2025)
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2024)
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2024)
S-Agents: Self-organizing Agents in Open-ended Environments
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
IA-T2I: Internet-Augmented Text-to-Image Generation
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
von: Shen, William F., et al.
Veröffentlicht: (2026)
von: Shen, William F., et al.
Veröffentlicht: (2026)
A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis
von: Imajo, Kentaro, et al.
Veröffentlicht: (2025)
von: Imajo, Kentaro, et al.
Veröffentlicht: (2025)
Towards Open-ended Visual Quality Comparison
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
PyVision-RL: Forging Open Agentic Vision Models via RL
von: Zhao, Shitian, et al.
Veröffentlicht: (2026)
von: Zhao, Shitian, et al.
Veröffentlicht: (2026)
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
von: Li, Teng, et al.
Veröffentlicht: (2025)
von: Li, Teng, et al.
Veröffentlicht: (2025)
Yume-1.5: A Text-Controlled Interactive World Generation Model
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2025)
OpenDDI: A Comprehensive Benchmark for DDI Prediction
von: Jin, Xinmo, et al.
Veröffentlicht: (2026)
von: Jin, Xinmo, et al.
Veröffentlicht: (2026)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
Cooperative Open-ended Learning Framework for Zero-shot Coordination
von: Li, Yang, et al.
Veröffentlicht: (2023)
von: Li, Yang, et al.
Veröffentlicht: (2023)
ECG Identity Authentication in Open-set with Multi-model Pretraining and Self-constraint Center & Irrelevant Sample Repulsion Learning
von: Dong, Mingyu, et al.
Veröffentlicht: (2025)
von: Dong, Mingyu, et al.
Veröffentlicht: (2025)
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
von: Chow, Wei, et al.
Veröffentlicht: (2025)
von: Chow, Wei, et al.
Veröffentlicht: (2025)
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
von: Lin, Yuqi, et al.
Veröffentlicht: (2025)
von: Lin, Yuqi, et al.
Veröffentlicht: (2025)
OpenSDI: Spotting Diffusion-Generated Images in the Open World
von: Wang, Yabin, et al.
Veröffentlicht: (2025)
von: Wang, Yabin, et al.
Veröffentlicht: (2025)
OpenGT: A Comprehensive Benchmark For Graph Transformers
von: Tang, Jiachen, et al.
Veröffentlicht: (2025)
von: Tang, Jiachen, et al.
Veröffentlicht: (2025)
Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory Matching
von: Guo, Ziyao, et al.
Veröffentlicht: (2023)
von: Guo, Ziyao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024) -
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025) -
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
von: Xu, Zhaopan, et al.
Veröffentlicht: (2025) -
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
von: Ai, Jiaxin, et al.
Veröffentlicht: (2025) -
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2024)