CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Dongzhi, Song, Guanglu, Wu, Xiaoshi, Zhang, Renrui, Shen, Dazhong, Zong, Zhuofan, Liu, Yu, Li, Hongsheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
di: Zong, Zhuofan, et al.
Pubblicazione: (2024)
di: Zong, Zhuofan, et al.
Pubblicazione: (2024)
MoVA: Adapting Mixture of Vision Experts to Multimodal Context
di: Zong, Zhuofan, et al.
Pubblicazione: (2024)
di: Zong, Zhuofan, et al.
Pubblicazione: (2024)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
ADT: Tuning Diffusion Models with Adversarial Supervision
di: Shen, Dazhong, et al.
Pubblicazione: (2025)
di: Shen, Dazhong, et al.
Pubblicazione: (2025)
Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models
di: Wu, Xiaoshi, et al.
Pubblicazione: (2024)
di: Wu, Xiaoshi, et al.
Pubblicazione: (2024)
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
di: Xue, Zeyue, et al.
Pubblicazione: (2023)
di: Xue, Zeyue, et al.
Pubblicazione: (2023)
Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
di: Ma, Bingqi, et al.
Pubblicazione: (2024)
di: Ma, Bingqi, et al.
Pubblicazione: (2024)
Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation
di: Wang, Fu-Yun, et al.
Pubblicazione: (2024)
di: Wang, Fu-Yun, et al.
Pubblicazione: (2024)
ECNet: Effective Controllable Text-to-Image Diffusion Models
di: Li, Sicheng, et al.
Pubblicazione: (2024)
di: Li, Sicheng, et al.
Pubblicazione: (2024)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
di: Huang, Jia-Hong, et al.
Pubblicazione: (2024)
di: Huang, Jia-Hong, et al.
Pubblicazione: (2024)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
di: Patel, Maitreya, et al.
Pubblicazione: (2023)
di: Patel, Maitreya, et al.
Pubblicazione: (2023)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
di: Lei, Jiayi, et al.
Pubblicazione: (2025)
di: Lei, Jiayi, et al.
Pubblicazione: (2025)
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models
di: Ruan, Chenxi, et al.
Pubblicazione: (2026)
di: Ruan, Chenxi, et al.
Pubblicazione: (2026)
ComCLIP: Training-Free Compositional Image and Text Matching
di: Jiang, Kenan, et al.
Pubblicazione: (2022)
di: Jiang, Kenan, et al.
Pubblicazione: (2022)
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models
di: Xue, Yuyang, et al.
Pubblicazione: (2025)
di: Xue, Yuyang, et al.
Pubblicazione: (2025)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
di: Shao, Hao, et al.
Pubblicazione: (2024)
di: Shao, Hao, et al.
Pubblicazione: (2024)
Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
di: Wang, Bin, et al.
Pubblicazione: (2024)
di: Wang, Bin, et al.
Pubblicazione: (2024)
Scaling Concept With Text-Guided Diffusion Models
di: Huang, Chao, et al.
Pubblicazione: (2024)
di: Huang, Chao, et al.
Pubblicazione: (2024)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
PALP: Prompt Aligned Personalization of Text-to-Image Models
di: Arar, Moab, et al.
Pubblicazione: (2024)
di: Arar, Moab, et al.
Pubblicazione: (2024)
GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
di: Liu, Zhaochen, et al.
Pubblicazione: (2026)
di: Liu, Zhaochen, et al.
Pubblicazione: (2026)
Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
di: Chin, Zhi-Yi, et al.
Pubblicazione: (2023)
di: Chin, Zhi-Yi, et al.
Pubblicazione: (2023)
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
di: Cui, Wanqing, et al.
Pubblicazione: (2024)
di: Cui, Wanqing, et al.
Pubblicazione: (2024)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
di: Tong, Chengzhuo, et al.
Pubblicazione: (2025)
di: Tong, Chengzhuo, et al.
Pubblicazione: (2025)
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
di: Wang, Jinyin, et al.
Pubblicazione: (2024)
di: Wang, Jinyin, et al.
Pubblicazione: (2024)
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
di: Motamed, Saman, et al.
Pubblicazione: (2023)
di: Motamed, Saman, et al.
Pubblicazione: (2023)
Match & Choose: Model Selection Framework for Fine-tuning Text-to-Image Diffusion Models
di: Lewandowski, Basile, et al.
Pubblicazione: (2025)
di: Lewandowski, Basile, et al.
Pubblicazione: (2025)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
di: Ye, Junyan, et al.
Pubblicazione: (2025)
di: Ye, Junyan, et al.
Pubblicazione: (2025)
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
di: Beňová, Ivana, et al.
Pubblicazione: (2024)
di: Beňová, Ivana, et al.
Pubblicazione: (2024)
Rethinking the Spatial Inconsistency in Classifier-Free Diffusion Guidance
di: Shen, Dazhong, et al.
Pubblicazione: (2024)
di: Shen, Dazhong, et al.
Pubblicazione: (2024)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
di: Lv, Zheqi, et al.
Pubblicazione: (2025)
di: Lv, Zheqi, et al.
Pubblicazione: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
di: Chen, Xinyan, et al.
Pubblicazione: (2025)
di: Chen, Xinyan, et al.
Pubblicazione: (2025)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
di: Tian, Changyao, et al.
Pubblicazione: (2024)
di: Tian, Changyao, et al.
Pubblicazione: (2024)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
di: Tan, Zhiyu, et al.
Pubblicazione: (2024)
di: Tan, Zhiyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025) -
EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
di: Zong, Zhuofan, et al.
Pubblicazione: (2024) -
MoVA: Adapting Mixture of Vision Experts to Multimodal Context
di: Zong, Zhuofan, et al.
Pubblicazione: (2024) -
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025) -
ADT: Tuning Diffusion Models with Adversarial Supervision
di: Shen, Dazhong, et al.
Pubblicazione: (2025)