CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Wei, Li, Lin, Yang, Yongqi, Wen, Bin, Yang, Fan, Gao, Tingting, Wu, Yu, Chen, Long |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
von: Chen, Wei, et al.
Veröffentlicht: (2026)
von: Chen, Wei, et al.
Veröffentlicht: (2026)
Towards Text-Image Interleaved Retrieval
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
CI-VID: A Coherent Interleaved Text-Video Dataset
von: Ju, Yiming, et al.
Veröffentlicht: (2025)
von: Ju, Yiming, et al.
Veröffentlicht: (2025)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
CoMM: Collaborative Multi-Agent, Multi-Reasoning-Path Prompting for Complex Problem Solving
von: Chen, Pei, et al.
Veröffentlicht: (2024)
von: Chen, Pei, et al.
Veröffentlicht: (2024)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
von: Li, Qingyun, et al.
Veröffentlicht: (2024)
von: Li, Qingyun, et al.
Veröffentlicht: (2024)
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
von: Luan, Bozhi, et al.
Veröffentlicht: (2024)
von: Luan, Bozhi, et al.
Veröffentlicht: (2024)
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
von: Wang, Bingli, et al.
Veröffentlicht: (2026)
von: Wang, Bingli, et al.
Veröffentlicht: (2026)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
von: Chen, Chao, et al.
Veröffentlicht: (2025)
von: Chen, Chao, et al.
Veröffentlicht: (2025)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
von: Feng, Yukang, et al.
Veröffentlicht: (2025)
von: Feng, Yukang, et al.
Veröffentlicht: (2025)
Interleaving Reasoning for Better Text-to-Image Generation
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
MM-ACT: Learn from Multimodal Parallel Generation to Act
von: Liang, Haotian, et al.
Veröffentlicht: (2025)
von: Liang, Haotian, et al.
Veröffentlicht: (2025)
MULTI: Multimodal Understanding Leaderboard with Text and Images
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
Bringing The Consistency Gap: Explicit Structured Memory for Interleaved Image-Text Generation
von: Lin, Zeteng, et al.
Veröffentlicht: (2025)
von: Lin, Zeteng, et al.
Veröffentlicht: (2025)
Holistic Evaluation for Interleaved Text-and-Image Generation
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
von: Wang, Junjie, et al.
Veröffentlicht: (2024)
von: Wang, Junjie, et al.
Veröffentlicht: (2024)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2023)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2023)
Interleaved-Modal Chain-of-Thought
von: Gao, Jun, et al.
Veröffentlicht: (2024)
von: Gao, Jun, et al.
Veröffentlicht: (2024)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
von: Wang, Yi, et al.
Veröffentlicht: (2023)
von: Wang, Yi, et al.
Veröffentlicht: (2023)
Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
von: Zou, Zhentao, et al.
Veröffentlicht: (2025)
von: Zou, Zhentao, et al.
Veröffentlicht: (2025)
Learning Video Context as Interleaved Multimodal Sequences
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
MM-ISTS: Cooperating Irregularly Sampled Time Series Forecasting with Multimodal Vision-Text LLMs
von: Lei, Zhi, et al.
Veröffentlicht: (2026)
von: Lei, Zhi, et al.
Veröffentlicht: (2026)
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
von: Chern, Ethan, et al.
Veröffentlicht: (2024)
von: Chern, Ethan, et al.
Veröffentlicht: (2024)
DuoGen: Towards General Purpose Interleaved Multimodal Generation
von: Shi, Min, et al.
Veröffentlicht: (2026)
von: Shi, Min, et al.
Veröffentlicht: (2026)
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
von: Wu, Shengqiong, et al.
Veröffentlicht: (2026)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2026)
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
von: Zhou, Pengfei, et al.
Veröffentlicht: (2024)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2024)
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2025)
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2025)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation
von: Lu, Xiaoxin, et al.
Veröffentlicht: (2025)
von: Lu, Xiaoxin, et al.
Veröffentlicht: (2025)
Diffusion in Diffusion: Cyclic One-Way Diffusion for Text-Vision-Conditioned Generation
von: Wang, Ruoyu, et al.
Veröffentlicht: (2023)
von: Wang, Ruoyu, et al.
Veröffentlicht: (2023)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
von: Sun, Linzhuang, et al.
Veröffentlicht: (2025)
von: Sun, Linzhuang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024) -
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
von: Chen, Dongping, et al.
Veröffentlicht: (2024) -
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
von: Chen, Wei, et al.
Veröffentlicht: (2026) -
Towards Text-Image Interleaved Retrieval
von: Zhang, Xin, et al.
Veröffentlicht: (2025) -
CI-VID: A Coherent Interleaved Text-Video Dataset
von: Ju, Yiming, et al.
Veröffentlicht: (2025)