A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Yukang, Sun, Jianwen, Li, Chuanhao, Li, Zizhen, Ai, Jiaxin, Zhang, Fanrui, Chang, Yifan, Zhou, Sizhuo, Zhang, Shenglin, Dai, Yu, Zhang, Kaipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
Closing the Expression Gap in LLM Instructions via Socratic Questioning
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
World Craft: Agentic Framework to Create Visualizable Worlds via Text
von: Sun, Jianwen, et al.
Veröffentlicht: (2026)
von: Sun, Jianwen, et al.
Veröffentlicht: (2026)
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
von: Zhang, Fanrui, et al.
Veröffentlicht: (2025)
von: Zhang, Fanrui, et al.
Veröffentlicht: (2025)
IA-T2I: Internet-Augmented Text-to-Image Generation
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
von: Chang, Yifan, et al.
Veröffentlicht: (2025)
von: Chang, Yifan, et al.
Veröffentlicht: (2025)
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
von: Ai, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ai, Jiaxin, et al.
Veröffentlicht: (2025)
MeepleLM: A Virtual Playtester Simulating Diverse Subjective Experiences
von: Li, Zizhen, et al.
Veröffentlicht: (2026)
von: Li, Zizhen, et al.
Veröffentlicht: (2026)
InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
von: Li, Zizhen, et al.
Veröffentlicht: (2025)
von: Li, Zizhen, et al.
Veröffentlicht: (2025)
AutoBG: A Board Game Design Assistant with Interactive Ideation, Iterative Rulebook Generation, and Individualized Feedback
von: Li, Zizhen, et al.
Veröffentlicht: (2026)
von: Li, Zizhen, et al.
Veröffentlicht: (2026)
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
von: Ai, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ai, Jiaxin, et al.
Veröffentlicht: (2025)
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
von: Feng, Yukang, et al.
Veröffentlicht: (2026)
von: Feng, Yukang, et al.
Veröffentlicht: (2026)
Sekai: A Video Dataset towards World Exploration
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
LaGen: Towards Autoregressive LiDAR Scene Generation
von: Zhou, Sizhuo, et al.
Veröffentlicht: (2025)
von: Zhou, Sizhuo, et al.
Veröffentlicht: (2025)
Holistic Evaluation for Interleaved Text-and-Image Generation
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
Towards Text-Image Interleaved Retrieval
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
von: Zhou, Pengfei, et al.
Veröffentlicht: (2024)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2024)
Global Context Compression with Interleaved Vision-Text Transformation
von: Jiao, Dian, et al.
Veröffentlicht: (2026)
von: Jiao, Dian, et al.
Veröffentlicht: (2026)
SVBench: Evaluation of Video Generation Models on Social Reasoning
von: Peng, Wenshuo, et al.
Veröffentlicht: (2025)
von: Peng, Wenshuo, et al.
Veröffentlicht: (2025)
WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG
von: Li, Zhen, et al.
Veröffentlicht: (2026)
von: Li, Zhen, et al.
Veröffentlicht: (2026)
Yume-1.5: A Text-Controlled Interactive World Generation Model
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2025)
The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation
von: Ruan, Shouwei, et al.
Veröffentlicht: (2025)
von: Ruan, Shouwei, et al.
Veröffentlicht: (2025)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
von: Fan, Cunxin, et al.
Veröffentlicht: (2025)
von: Fan, Cunxin, et al.
Veröffentlicht: (2025)
CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability,Reproducibility, and Practicality
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
von: Zhao, Shitian, et al.
Veröffentlicht: (2025)
von: Zhao, Shitian, et al.
Veröffentlicht: (2025)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
CI-VID: A Coherent Interleaved Text-Video Dataset
von: Ju, Yiming, et al.
Veröffentlicht: (2025)
von: Ju, Yiming, et al.
Veröffentlicht: (2025)
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
von: Liu, Jinkun, et al.
Veröffentlicht: (2026)
von: Liu, Jinkun, et al.
Veröffentlicht: (2026)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2023)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2023)
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2026)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
T3M: Text Guided 3D Human Motion Synthesis from Speech
von: Peng, Wenshuo, et al.
Veröffentlicht: (2024)
von: Peng, Wenshuo, et al.
Veröffentlicht: (2024)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
15M Multimodal Facial Image-Text Dataset
von: Dai, Dawei, et al.
Veröffentlicht: (2024)
von: Dai, Dawei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
von: Sun, Jianwen, et al.
Veröffentlicht: (2025) -
Closing the Expression Gap in LLM Instructions via Socratic Questioning
von: Sun, Jianwen, et al.
Veröffentlicht: (2025) -
From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
von: Sun, Jianwen, et al.
Veröffentlicht: (2025) -
World Craft: Agentic Framework to Create Visualizable Worlds via Text
von: Sun, Jianwen, et al.
Veröffentlicht: (2026) -
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
von: Zhang, Fanrui, et al.
Veröffentlicht: (2025)