DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Dongzhi, Zhang, Renrui, Li, Haodong, Zong, Zhuofan, Guo, Ziyu, He, Jun, Guo, Claire, Ye, Junyan, Fang, Rongyao, Li, Weijia, Liu, Rui, Li, Hongsheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
di: Li, Haodong, et al.
Pubblicazione: (2026)
di: Li, Haodong, et al.
Pubblicazione: (2026)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
di: Chen, Xinyan, et al.
Pubblicazione: (2025)
di: Chen, Xinyan, et al.
Pubblicazione: (2025)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
di: Tong, Chengzhuo, et al.
Pubblicazione: (2025)
di: Tong, Chengzhuo, et al.
Pubblicazione: (2025)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
di: Shao, Hao, et al.
Pubblicazione: (2024)
di: Shao, Hao, et al.
Pubblicazione: (2024)
FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
di: Zhu, Leqi, et al.
Pubblicazione: (2026)
di: Zhu, Leqi, et al.
Pubblicazione: (2026)
CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
di: Duan, Chengqi, et al.
Pubblicazione: (2025)
di: Duan, Chengqi, et al.
Pubblicazione: (2025)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
di: Gao, Peng, et al.
Pubblicazione: (2021)
di: Gao, Peng, et al.
Pubblicazione: (2021)
The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning
di: Zheng, Tianshi, et al.
Pubblicazione: (2025)
di: Zheng, Tianshi, et al.
Pubblicazione: (2025)
EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
di: Zong, Zhuofan, et al.
Pubblicazione: (2024)
di: Zong, Zhuofan, et al.
Pubblicazione: (2024)
MoVA: Adapting Mixture of Vision Experts to Multimodal Context
di: Zong, Zhuofan, et al.
Pubblicazione: (2024)
di: Zong, Zhuofan, et al.
Pubblicazione: (2024)
Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation
di: He, Jun, et al.
Pubblicazione: (2026)
di: He, Jun, et al.
Pubblicazione: (2026)
Automated Movie Generation via Multi-Agent CoT Planning
di: Wu, Weijia, et al.
Pubblicazione: (2025)
di: Wu, Weijia, et al.
Pubblicazione: (2025)
BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
di: Ye, Junyan, et al.
Pubblicazione: (2025)
di: Ye, Junyan, et al.
Pubblicazione: (2025)
GenClaw: Code-Driven Agentic Image Generation
di: Ye, Junyan, et al.
Pubblicazione: (2026)
di: Ye, Junyan, et al.
Pubblicazione: (2026)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
di: Deng, Yuntian, et al.
Pubblicazione: (2024)
di: Deng, Yuntian, et al.
Pubblicazione: (2024)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
di: Zhou, Weibo, et al.
Pubblicazione: (2025)
di: Zhou, Weibo, et al.
Pubblicazione: (2025)
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
di: Ye, Junyan, et al.
Pubblicazione: (2025)
di: Ye, Junyan, et al.
Pubblicazione: (2025)
CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs
di: Li, Li, et al.
Pubblicazione: (2025)
di: Li, Li, et al.
Pubblicazione: (2025)
Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
di: Ma, Bingqi, et al.
Pubblicazione: (2024)
di: Ma, Bingqi, et al.
Pubblicazione: (2024)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
di: Liao, Jiaqi, et al.
Pubblicazione: (2025)
di: Liao, Jiaqi, et al.
Pubblicazione: (2025)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
di: Li, Ang, et al.
Pubblicazione: (2025)
di: Li, Ang, et al.
Pubblicazione: (2025)
CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
di: Wang, Guankun, et al.
Pubblicazione: (2024)
di: Wang, Guankun, et al.
Pubblicazione: (2024)
ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
di: Wang, Lihong, et al.
Pubblicazione: (2025)
di: Wang, Lihong, et al.
Pubblicazione: (2025)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
di: Ren, Ruifeng, et al.
Pubblicazione: (2024)
di: Ren, Ruifeng, et al.
Pubblicazione: (2024)
MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
di: Sprague, Zayne, et al.
Pubblicazione: (2024)
di: Sprague, Zayne, et al.
Pubblicazione: (2024)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
di: Ye, Junyan, et al.
Pubblicazione: (2025)
di: Ye, Junyan, et al.
Pubblicazione: (2025)
Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
di: Jin, Senjie, et al.
Pubblicazione: (2025)
di: Jin, Senjie, et al.
Pubblicazione: (2025)
LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA
di: Huang, Jing, et al.
Pubblicazione: (2025)
di: Huang, Jing, et al.
Pubblicazione: (2025)
Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning
di: Shen, Zijun, et al.
Pubblicazione: (2026)
di: Shen, Zijun, et al.
Pubblicazione: (2026)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
di: Lei, Jiayi, et al.
Pubblicazione: (2025)
di: Lei, Jiayi, et al.
Pubblicazione: (2025)
Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
di: Li, Jiatong, et al.
Pubblicazione: (2025)
di: Li, Jiatong, et al.
Pubblicazione: (2025)
Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
di: Le, Chenqian, et al.
Pubblicazione: (2025)
di: Le, Chenqian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025) -
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024) -
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
di: Li, Haodong, et al.
Pubblicazione: (2026) -
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025) -
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
di: Chen, Xinyan, et al.
Pubblicazione: (2025)