Reflective Human-Machine Co-adaptation for Enhanced Text-to-Image Generation Dialogue System
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yuheng, He, Yangfan, Xia, Yinghui, Shi, Tianyu, Wang, Jun, Yang, Jinsong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Twin Co-Adaptive Dialogue for Progressive Image Generation
by: Wang, Jianhui, et al.
Published: (2025)
by: Wang, Jianhui, et al.
Published: (2025)
TDRI: Two-Phase Dialogue Refinement and Co-Adaptation for Interactive Image Generation
by: Feng, Yuheng, et al.
Published: (2025)
by: Feng, Yuheng, et al.
Published: (2025)
DDPM-MoCo: Advancing Industrial Surface Defect Generation and Detection with Generative and Contrastive Learning
by: He, Yangfan, et al.
Published: (2024)
by: He, Yangfan, et al.
Published: (2024)
PurifyGen: A Risk-Discrimination and Semantic-Purification Model for Safe Text-to-Image Generation
by: Cao, Zongsheng, et al.
Published: (2025)
by: Cao, Zongsheng, et al.
Published: (2025)
Enhancing Intent Understanding for Ambiguous prompt: A Human-Machine Co-Adaption Strategy
by: He, Yangfan, et al.
Published: (2025)
by: He, Yangfan, et al.
Published: (2025)
MARS: Memory-Enhanced Agents with Reflective Self-improvement
by: Liang, Xuechen, et al.
Published: (2025)
by: Liang, Xuechen, et al.
Published: (2025)
Rich Human Feedback for Text-to-Image Generation
by: Liang, Youwei, et al.
Published: (2023)
by: Liang, Youwei, et al.
Published: (2023)
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
by: Luan, Bozhi, et al.
Published: (2024)
by: Luan, Bozhi, et al.
Published: (2024)
AnatoMask: Enhancing Medical Image Segmentation with Reconstruction-guided Self-masking
by: Li, Yuheng, et al.
Published: (2024)
by: Li, Yuheng, et al.
Published: (2024)
CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
by: Cao, Zongsheng, et al.
Published: (2025)
by: Cao, Zongsheng, et al.
Published: (2025)
Free-Mask: A Novel Paradigm of Integration Between the Segmentation Diffusion Model and Image Editing
by: Gao, Bo, et al.
Published: (2024)
by: Gao, Bo, et al.
Published: (2024)
DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
by: Huang, Minbin, et al.
Published: (2024)
by: Huang, Minbin, et al.
Published: (2024)
WcDT: World-centric Diffusion Transformer for Traffic Scene Generation
by: Yang, Chen, et al.
Published: (2024)
by: Yang, Chen, et al.
Published: (2024)
Text2Traffic: A Text-to-Image Generation and Editing Method for Traffic Scenes
by: Lv, Feng, et al.
Published: (2025)
by: Lv, Feng, et al.
Published: (2025)
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
by: Ba, Ying, et al.
Published: (2025)
by: Ba, Ying, et al.
Published: (2025)
CoMo: Compositional Motion Customization for Text-to-Video Generation
by: Xu, Youcan, et al.
Published: (2025)
by: Xu, Youcan, et al.
Published: (2025)
Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation
by: Li, Yaqi, et al.
Published: (2025)
by: Li, Yaqi, et al.
Published: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
by: Yang, Lehan, et al.
Published: (2025)
by: Yang, Lehan, et al.
Published: (2025)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
by: Liao, Jiaqi, et al.
Published: (2025)
by: Liao, Jiaqi, et al.
Published: (2025)
FASIONAD : FAst and Slow FusION Thinking Systems for Human-Like Autonomous Driving with Adaptive Feedback
by: Qian, Kangan, et al.
Published: (2024)
by: Qian, Kangan, et al.
Published: (2024)
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
by: Yue, Yang, et al.
Published: (2026)
by: Yue, Yang, et al.
Published: (2026)
TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt
by: Wu, Xiangyu, et al.
Published: (2024)
by: Wu, Xiangyu, et al.
Published: (2024)
CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation
by: Gao, Zhanxin, et al.
Published: (2025)
by: Gao, Zhanxin, et al.
Published: (2025)
EmoGen: Emotional Image Content Generation with Text-to-Image Diffusion Models
by: Yang, Jingyuan, et al.
Published: (2024)
by: Yang, Jingyuan, et al.
Published: (2024)
DeCoT: Decomposing Complex Instructions for Enhanced Text-to-Image Generation with Large Language Models
by: Lin, Xiaochuan, et al.
Published: (2025)
by: Lin, Xiaochuan, et al.
Published: (2025)
Flatten: Video Action Recognition is an Image Classification task
by: Chen, Junlin, et al.
Published: (2024)
by: Chen, Junlin, et al.
Published: (2024)
Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation
by: He, Yiguo, et al.
Published: (2025)
by: He, Yiguo, et al.
Published: (2025)
Dynamic Prompt Optimizing for Text-to-Image Generation
by: Mo, Wenyi, et al.
Published: (2024)
by: Mo, Wenyi, et al.
Published: (2024)
Unified Prompt Attack Against Text-to-Image Generation Models
by: Peng, Duo, et al.
Published: (2025)
by: Peng, Duo, et al.
Published: (2025)
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
by: Qian, Zhaofang, et al.
Published: (2024)
by: Qian, Zhaofang, et al.
Published: (2024)
Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative Prior
by: Feng, Ruoyu, et al.
Published: (2025)
by: Feng, Ruoyu, et al.
Published: (2025)
Controllable Generation with Text-to-Image Diffusion Models: A Survey
by: Cao, Pu, et al.
Published: (2024)
by: Cao, Pu, et al.
Published: (2024)
TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
by: Cao, Zongsheng, et al.
Published: (2025)
by: Cao, Zongsheng, et al.
Published: (2025)
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
by: Zhao, Chengfeng, et al.
Published: (2026)
by: Zhao, Chengfeng, et al.
Published: (2026)
Iterative Prompt Refinement for Safer Text-to-Image Generation
by: Jeon, Jinwoo, et al.
Published: (2025)
by: Jeon, Jinwoo, et al.
Published: (2025)
Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Preference Understanding
by: Li, Kun, et al.
Published: (2025)
by: Li, Kun, et al.
Published: (2025)
MobiDiary: Autoregressive Action Captioning with Wearable Devices and Wireless Signals
by: Deng, Fei, et al.
Published: (2026)
by: Deng, Fei, et al.
Published: (2026)
Similar Items
-
Twin Co-Adaptive Dialogue for Progressive Image Generation
by: Wang, Jianhui, et al.
Published: (2025) -
TDRI: Two-Phase Dialogue Refinement and Co-Adaptation for Interactive Image Generation
by: Feng, Yuheng, et al.
Published: (2025) -
DDPM-MoCo: Advancing Industrial Surface Defect Generation and Detection with Generative and Contrastive Learning
by: He, Yangfan, et al.
Published: (2024) -
PurifyGen: A Risk-Discrimination and Semantic-Purification Model for Safe Text-to-Image Generation
by: Cao, Zongsheng, et al.
Published: (2025) -
Enhancing Intent Understanding for Ambiguous prompt: A Human-Machine Co-Adaption Strategy
by: He, Yangfan, et al.
Published: (2025)