BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Dewei, Li, Mingwei, Yang, Zongxin, Lu, Yu, Xu, Yunqiu, Wang, Zhizhong, Huang, Zeyi, Yang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
by: Zhou, Dewei, et al.
Published: (2025)
by: Zhou, Dewei, et al.
Published: (2025)
3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering
by: Zhou, Dewei, et al.
Published: (2025)
by: Zhou, Dewei, et al.
Published: (2025)
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
by: Zhou, Dewei, et al.
Published: (2026)
by: Zhou, Dewei, et al.
Published: (2026)
SIFU: Side-view Conditioned Implicit Function for Real-world Usable Clothed Human Reconstruction
by: Zhang, Zechuan, et al.
Published: (2023)
by: Zhang, Zechuan, et al.
Published: (2023)
Controllable 3D Face Generation with Conditional Style Code Diffusion
by: Shen, Xiaolong, et al.
Published: (2023)
by: Shen, Xiaolong, et al.
Published: (2023)
DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
by: Wang, Zhizhong, et al.
Published: (2025)
by: Wang, Zhizhong, et al.
Published: (2025)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
by: Zhang, Zechuan, et al.
Published: (2025)
by: Zhang, Zechuan, et al.
Published: (2025)
Text-Image Conditioned 3D Generation
by: Cen, Jiazhong, et al.
Published: (2026)
by: Cen, Jiazhong, et al.
Published: (2026)
Origin Identification for Text-Guided Image-to-Image Diffusion Models
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
3D Object Manipulation in a Single Image using Generative Models
by: Zhao, Ruisi, et al.
Published: (2025)
by: Zhao, Ruisi, et al.
Published: (2025)
MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
Conditional Text-to-Image Generation with Reference Guidance
by: Kim, Taewook, et al.
Published: (2024)
by: Kim, Taewook, et al.
Published: (2024)
HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
by: Zhou, Zhenglin, et al.
Published: (2024)
by: Zhou, Zhenglin, et al.
Published: (2024)
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
by: Xu, Yunqiu, et al.
Published: (2024)
by: Xu, Yunqiu, et al.
Published: (2024)
Are Image-to-Video Models Good Zero-Shot Image Editors?
by: Zhang, Zechuan, et al.
Published: (2025)
by: Zhang, Zechuan, et al.
Published: (2025)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
by: Peng, Bo, et al.
Published: (2023)
by: Peng, Bo, et al.
Published: (2023)
ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
by: Xu, Ruihang, et al.
Published: (2025)
by: Xu, Ruihang, et al.
Published: (2025)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
by: Zhao, Zixu, et al.
Published: (2025)
by: Zhao, Zixu, et al.
Published: (2025)
Fast Prompt Alignment for Text-to-Image Generation
by: Mrini, Khalil, et al.
Published: (2024)
by: Mrini, Khalil, et al.
Published: (2024)
Object-Conditioned Energy-Based Attention Map Alignment in Text-to-Image Diffusion Models
by: Zhang, Yasi, et al.
Published: (2024)
by: Zhang, Yasi, et al.
Published: (2024)
Joint Conditional Diffusion Model for Image Restoration with Mixed Degradations
by: Yue, Yufeng, et al.
Published: (2024)
by: Yue, Yufeng, et al.
Published: (2024)
Attack Deterministic Conditional Image Generative Models for Diverse and Controllable Generation
by: Chu, Tianyi, et al.
Published: (2024)
by: Chu, Tianyi, et al.
Published: (2024)
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
by: Huang, Qihan, et al.
Published: (2024)
by: Huang, Qihan, et al.
Published: (2024)
Local Conditional Controlling for Text-to-Image Diffusion Models
by: Zhao, Yibo, et al.
Published: (2023)
by: Zhao, Yibo, et al.
Published: (2023)
UDiFF: Generating Conditional Unsigned Distance Fields with Optimal Wavelet Diffusion
by: Zhou, Junsheng, et al.
Published: (2024)
by: Zhou, Junsheng, et al.
Published: (2024)
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
Image Inpainting via Conditional Texture and Structure Dual Generation
by: Guo, Xiefan, et al.
Published: (2021)
by: Guo, Xiefan, et al.
Published: (2021)
GD^2-NeRF: Generative Detail Compensation via GAN and Diffusion for One-shot Generalizable Neural Radiance Fields
by: Pan, Xiao, et al.
Published: (2024)
by: Pan, Xiao, et al.
Published: (2024)
PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing
by: Xu, Ruihang, et al.
Published: (2026)
by: Xu, Ruihang, et al.
Published: (2026)
Simultaneous Enhancement and Noise Suppression under Complex Illumination Conditions
by: Tao, Jing, et al.
Published: (2025)
by: Tao, Jing, et al.
Published: (2025)
Language-Image Alignment with Fixed Text Encoders
by: Yang, Jingfeng, et al.
Published: (2025)
by: Yang, Jingfeng, et al.
Published: (2025)
Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
by: Koo, Jaywon, et al.
Published: (2025)
by: Koo, Jaywon, et al.
Published: (2025)
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
by: Zheng, Zirui, et al.
Published: (2025)
by: Zheng, Zirui, et al.
Published: (2025)
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Diffusion in Diffusion: Cyclic One-Way Diffusion for Text-Vision-Conditioned Generation
by: Wang, Ruoyu, et al.
Published: (2023)
by: Wang, Ruoyu, et al.
Published: (2023)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Progressive Image Restoration via Text-Conditioned Video Generation
by: Kang, Peng, et al.
Published: (2025)
by: Kang, Peng, et al.
Published: (2025)
Similar Items
-
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
by: Zhou, Dewei, et al.
Published: (2025) -
3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
by: Zhou, Dewei, et al.
Published: (2024) -
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024) -
3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering
by: Zhou, Dewei, et al.
Published: (2025) -
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
by: Zhou, Dewei, et al.
Published: (2026)