Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Yi, Liu, Mushui, He, Wanggui, Yuan, Hanyang, Zhang, Longxiang, Huang, Ziwei, Zhang, Guanghao, Fang, Wenkai, Jiang, Haoze, Zhang, Shengxuming, She, Dong, Liu, Jinlong, Dai, Weilong, Song, Mingli, Jiang, Hao, Song, Jie |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
par: Huang, Qihan, et autres
Publié: (2024)
par: Huang, Qihan, et autres
Publié: (2024)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
par: Zhang, Peng, et autres
Publié: (2026)
par: Zhang, Peng, et autres
Publié: (2026)
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
par: Zhang, Guanghao, et autres
Publié: (2025)
par: Zhang, Guanghao, et autres
Publié: (2025)
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
par: Zhang, Longxiang, et autres
Publié: (2026)
par: Zhang, Longxiang, et autres
Publié: (2026)
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
par: Liu, Jinlong, et autres
Publié: (2026)
par: Liu, Jinlong, et autres
Publié: (2026)
Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
par: Huang, Qihan, et autres
Publié: (2025)
par: Huang, Qihan, et autres
Publié: (2025)
UniMo: Unified Motion Generation and Understanding with Chain of Thought
par: Wang, Guocun, et autres
Publié: (2026)
par: Wang, Guocun, et autres
Publié: (2026)
Dataset Ownership Verification for Pre-trained Masked Models
par: Xie, Yuechen, et autres
Publié: (2025)
par: Xie, Yuechen, et autres
Publié: (2025)
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO
par: Tong, Yunze, et autres
Publié: (2026)
par: Tong, Yunze, et autres
Publié: (2026)
CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers
par: She, D., et autres
Publié: (2025)
par: She, D., et autres
Publié: (2025)
MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
par: She, Dong, et autres
Publié: (2025)
par: She, Dong, et autres
Publié: (2025)
Liquid: Language Models are Scalable and Unified Multi-modal Generators
par: Wu, Junfeng, et autres
Publié: (2024)
par: Wu, Junfeng, et autres
Publié: (2024)
Chain-of-Thought Poisoning Attacks against R1-based Retrieval-Augmented Generation Systems
par: Song, Hongru, et autres
Publié: (2025)
par: Song, Hongru, et autres
Publié: (2025)
Efficient and Comprehensive Feature Extraction in Large Vision-Language Model for Pathology Analysis
par: Zhang, Shengxuming, et autres
Publié: (2024)
par: Zhang, Shengxuming, et autres
Publié: (2024)
SCoTER: Structured Chain-of-Thought Transfer for Enhanced Recommendation
par: Jiang, Jie, et autres
Publié: (2025)
par: Jiang, Jie, et autres
Publié: (2025)
A phenomenology condition other than zero resistance and a possible pairing mechanism of holes-electrons for superconductivity
par: She, Weilong
Publié: (2011)
par: She, Weilong
Publié: (2011)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
par: Gu, Chenyang, et autres
Publié: (2025)
par: Gu, Chenyang, et autres
Publié: (2025)
Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
par: Jiang, Jingjing, et autres
Publié: (2025)
par: Jiang, Jingjing, et autres
Publié: (2025)
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
par: Jiang, Gangwei, et autres
Publié: (2025)
par: Jiang, Gangwei, et autres
Publié: (2025)
DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
par: Wu, Fangtai, et autres
Publié: (2025)
par: Wu, Fangtai, et autres
Publié: (2025)
Deep learning‐based accurate diagnosis and quantitative evaluation of microvascular invasion in hepatocellular carcinoma on whole‐slide histopathology images
par: Xiuming Zhang, et autres
Publié: (2024)
par: Xiuming Zhang, et autres
Publié: (2024)
Physics-informed Diffusion Generation for Geomagnetic Map Interpolation
par: Li, Wenda, et autres
Publié: (2026)
par: Li, Wenda, et autres
Publié: (2026)
Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
par: Lucas, Ryan, et autres
Publié: (2025)
par: Lucas, Ryan, et autres
Publié: (2025)
Hybrid Mask Generation for Infrared Small Target Detection with Single-Point Supervision
par: He, Weijie, et autres
Publié: (2024)
par: He, Weijie, et autres
Publié: (2024)
Unified Personalized Understanding, Generating and Editing
par: Zhong, Yu, et autres
Publié: (2026)
par: Zhong, Yu, et autres
Publié: (2026)
Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
par: Lam, Max W. Y., et autres
Publié: (2025)
par: Lam, Max W. Y., et autres
Publié: (2025)
Upfront Chain-of-Thought: A Cooperative Framework for Chain-of-Thought Compression
par: Li, Chengzhengxu, et autres
Publié: (2025)
par: Li, Chengzhengxu, et autres
Publié: (2025)
Reasoning with Reinforced Functional Token Tuning
par: Zhang, Kongcheng, et autres
Publié: (2025)
par: Zhang, Kongcheng, et autres
Publié: (2025)
Odyssey: Empowering Minecraft Agents with Open-World Skills
par: Liu, Shunyu, et autres
Publié: (2024)
par: Liu, Shunyu, et autres
Publié: (2024)
SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data
par: Fang, Wenkai, et autres
Publié: (2025)
par: Fang, Wenkai, et autres
Publié: (2025)
Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
par: Xu, Haolei, et autres
Publié: (2025)
par: Xu, Haolei, et autres
Publié: (2025)
MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis
par: He, Wanggui, et autres
Publié: (2024)
par: He, Wanggui, et autres
Publié: (2024)
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
par: Wang, Yaoting, et autres
Publié: (2025)
par: Wang, Yaoting, et autres
Publié: (2025)
SEW: Self-calibration Enhanced Whole Slide Pathology Image Analysis
par: Luo, Haoming, et autres
Publié: (2024)
par: Luo, Haoming, et autres
Publié: (2024)
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
par: Chen, Zihao, et autres
Publié: (2024)
par: Chen, Zihao, et autres
Publié: (2024)
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
par: Song, Lin, et autres
Publié: (2026)
par: Song, Lin, et autres
Publié: (2026)
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
par: Wu, Linquan, et autres
Publié: (2026)
par: Wu, Linquan, et autres
Publié: (2026)
Enhancing Generalization in Chain of Thought Reasoning for Smaller Models
par: Yin, Maxwell J., et autres
Publié: (2025)
par: Yin, Maxwell J., et autres
Publié: (2025)
GEWUM: General Exploration Workflow for the Utopia of Materials: A Unified Platform for Automated Structure Generation, Selection, and Validation
par: Song, Jiexi, et autres
Publié: (2026)
par: Song, Jiexi, et autres
Publié: (2026)
On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning
par: Li, Zhaoyi, et autres
Publié: (2026)
par: Li, Zhaoyi, et autres
Publié: (2026)
Documents similaires
-
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
par: Huang, Qihan, et autres
Publié: (2024) -
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
par: Zhang, Peng, et autres
Publié: (2026) -
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
par: Zhang, Guanghao, et autres
Publié: (2025) -
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
par: Zhang, Longxiang, et autres
Publié: (2026) -
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
par: Liu, Jinlong, et autres
Publié: (2026)