M3: High-fidelity Text-to-Image Generation via Multi-Modal, Multi-Agent and Multi-Round Visual Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Bangji, Guo, Ruihan, Fan, Jiajun, Cheng, Chaoran, Liu, Ge |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical Text
by: Xu, Ronghao, et al.
Published: (2025)
by: Xu, Ronghao, et al.
Published: (2025)
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
by: Sun, Wenfang, et al.
Published: (2026)
by: Sun, Wenfang, et al.
Published: (2026)
M3-AGIQA: Multimodal, Multi-Round, Multi-Aspect AI-Generated Image Quality Assessment
by: Cui, Chuan, et al.
Published: (2025)
by: Cui, Chuan, et al.
Published: (2025)
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
by: Tong, Qinyue, et al.
Published: (2025)
by: Tong, Qinyue, et al.
Published: (2025)
High-fidelity Multi-view Normal Integration with Scale-encoded Neural Surface Representation
by: Yang, Tongyu, et al.
Published: (2026)
by: Yang, Tongyu, et al.
Published: (2026)
PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images
by: Zhang, Shuoshuo, et al.
Published: (2025)
by: Zhang, Shuoshuo, et al.
Published: (2025)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
by: Ding, Yanbo, et al.
Published: (2024)
by: Ding, Yanbo, et al.
Published: (2024)
CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder
by: Ma, Lichen, et al.
Published: (2024)
by: Ma, Lichen, et al.
Published: (2024)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing
by: Ma, Shichao, et al.
Published: (2025)
by: Ma, Shichao, et al.
Published: (2025)
Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution
by: Wang, Ying, et al.
Published: (2023)
by: Wang, Ying, et al.
Published: (2023)
MMGR: Multi-Modal Generative Reasoning
by: Cai, Zefan, et al.
Published: (2025)
by: Cai, Zefan, et al.
Published: (2025)
Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization
by: Wang, Xiaohan, et al.
Published: (2025)
by: Wang, Xiaohan, et al.
Published: (2025)
Cocktail: Mixing Multi-Modality Controls for Text-Conditional Image Generation
by: Hu, Minghui, et al.
Published: (2023)
by: Hu, Minghui, et al.
Published: (2023)
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
by: Zhang, Haochen, et al.
Published: (2026)
by: Zhang, Haochen, et al.
Published: (2026)
Lightweight Method for Interactive 3D Medical Image Segmentation with Multi-Round Result Fusion
by: Shen, Bingzhi, et al.
Published: (2024)
by: Shen, Bingzhi, et al.
Published: (2024)
WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
by: Guan, Runwei, et al.
Published: (2026)
by: Guan, Runwei, et al.
Published: (2026)
ChartM$^3$: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehension
by: Xu, Duo, et al.
Published: (2025)
by: Xu, Duo, et al.
Published: (2025)
CMMD: Contrastive Multi-Modal Diffusion for Video-Audio Conditional Modeling
by: Yang, Ruihan, et al.
Published: (2023)
by: Yang, Ruihan, et al.
Published: (2023)
Can Text-to-image Model Assist Multi-modal Learning for Visual Recognition with Visual Modality Missing?
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning
by: Wang, Alex Jinpeng, et al.
Published: (2024)
by: Wang, Alex Jinpeng, et al.
Published: (2024)
MaxFusion: Plug&Play Multi-Modal Generation in Text-to-Image Diffusion Models
by: Nair, Nithin Gopalakrishnan, et al.
Published: (2024)
by: Nair, Nithin Gopalakrishnan, et al.
Published: (2024)
Multi-Modal 3D Mesh Reconstruction from Images and Text
by: Reka, Melvin, et al.
Published: (2025)
by: Reka, Melvin, et al.
Published: (2025)
TF-TI2I: Training-Free Text-and-Image-to-Image Generation via Multi-Modal Implicit-Context Learning in Text-to-Image Models
by: Hsiao, Teng-Fang, et al.
Published: (2025)
by: Hsiao, Teng-Fang, et al.
Published: (2025)
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
CustomTex: High-fidelity Indoor Scene Texturing via Multi-Reference Customization
by: Chen, Weilin, et al.
Published: (2026)
by: Chen, Weilin, et al.
Published: (2026)
Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld
by: Yang, Yijun, et al.
Published: (2023)
by: Yang, Yijun, et al.
Published: (2023)
MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
by: Li, Mingcheng, et al.
Published: (2025)
by: Li, Mingcheng, et al.
Published: (2025)
M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
by: Bai, Fan, et al.
Published: (2024)
by: Bai, Fan, et al.
Published: (2024)
Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization
by: Li, Jinghan, et al.
Published: (2026)
by: Li, Jinghan, et al.
Published: (2026)
MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation
by: Zhang, Xujie, et al.
Published: (2024)
by: Zhang, Xujie, et al.
Published: (2024)
PoIFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest
by: Deng, Jiajun, et al.
Published: (2024)
by: Deng, Jiajun, et al.
Published: (2024)
M4V: Multi-Modal Mamba for Text-to-Video Generation
by: Huang, Jiancheng, et al.
Published: (2025)
by: Huang, Jiancheng, et al.
Published: (2025)
EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts
by: Han, Yucheng, et al.
Published: (2024)
by: Han, Yucheng, et al.
Published: (2024)
Collaborative Multi-Modal Coding for High-Quality 3D Generation
by: Cao, Ziang, et al.
Published: (2025)
by: Cao, Ziang, et al.
Published: (2025)
RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
by: Hu, Huiyang, et al.
Published: (2025)
by: Hu, Huiyang, et al.
Published: (2025)
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
by: Li, Qian, et al.
Published: (2024)
by: Li, Qian, et al.
Published: (2024)
DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
by: Huang, Minbin, et al.
Published: (2024)
by: Huang, Minbin, et al.
Published: (2024)
Sculpt3D: Multi-View Consistent Text-to-3D Generation with Sparse 3D Prior
by: Chen, Cheng, et al.
Published: (2024)
by: Chen, Cheng, et al.
Published: (2024)
Similar Items
-
MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical Text
by: Xu, Ronghao, et al.
Published: (2025) -
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
by: Sun, Wenfang, et al.
Published: (2026) -
M3-AGIQA: Multimodal, Multi-Round, Multi-Aspect AI-Generated Image Quality Assessment
by: Cui, Chuan, et al.
Published: (2025) -
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
by: Tong, Qinyue, et al.
Published: (2025) -
High-fidelity Multi-view Normal Integration with Scale-encoded Neural Surface Representation
by: Yang, Tongyu, et al.
Published: (2026)