Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hahn, Meera, Zeng, Wenjun, Kannen, Nithish, Galt, Rich, Badola, Kartikeya, Kim, Been, Wang, Zi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Aesthetics: Cultural Competence in Text-to-Image Models
von: Kannen, Nithish, et al.
Veröffentlicht: (2024)
von: Kannen, Nithish, et al.
Veröffentlicht: (2024)
Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing
von: Ma, Shichao, et al.
Veröffentlicht: (2025)
von: Ma, Shichao, et al.
Veröffentlicht: (2025)
Alchemist: Turning Public Text-to-Image Data into Generative Gold
von: Startsev, Valerii, et al.
Veröffentlicht: (2025)
von: Startsev, Valerii, et al.
Veröffentlicht: (2025)
Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits
von: Kalibhat, Neha, et al.
Veröffentlicht: (2026)
von: Kalibhat, Neha, et al.
Veröffentlicht: (2026)
Instilling Multi-round Thinking to Text-guided Image Generation
von: Zeng, Lidong, et al.
Veröffentlicht: (2024)
von: Zeng, Lidong, et al.
Veröffentlicht: (2024)
Learning Complex Non-Rigid Image Edits from Multimodal Conditioning
von: Warner, Nikolai, et al.
Veröffentlicht: (2024)
von: Warner, Nikolai, et al.
Veröffentlicht: (2024)
UCMNet: Uncertainty-Aware Context Memory Network for Under-Display Camera Image Restoration
von: Kim, Daehyun, et al.
Veröffentlicht: (2026)
von: Kim, Daehyun, et al.
Veröffentlicht: (2026)
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
von: Li, Mingcheng, et al.
Veröffentlicht: (2025)
von: Li, Mingcheng, et al.
Veröffentlicht: (2025)
Generation Navigator: A State-Aware Agentic Framework for Image Generation
von: Liu, Jinming, et al.
Veröffentlicht: (2026)
von: Liu, Jinming, et al.
Veröffentlicht: (2026)
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
von: Chen, Yiyang, et al.
Veröffentlicht: (2022)
von: Chen, Yiyang, et al.
Veröffentlicht: (2022)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
von: Park, NaHyeon, et al.
Veröffentlicht: (2024)
von: Park, NaHyeon, et al.
Veröffentlicht: (2024)
Flow of Truth: Proactive Temporal Forensics for Image-to-Video Generation
von: Chen, Yuzhuo, et al.
Veröffentlicht: (2026)
von: Chen, Yuzhuo, et al.
Veröffentlicht: (2026)
ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection
von: Ma, Ruize, et al.
Veröffentlicht: (2025)
von: Ma, Ruize, et al.
Veröffentlicht: (2025)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
von: Zhao, Shihao, et al.
Veröffentlicht: (2024)
von: Zhao, Shihao, et al.
Veröffentlicht: (2024)
M3: High-fidelity Text-to-Image Generation via Multi-Modal, Multi-Agent and Multi-Round Visual Reasoning
von: Yang, Bangji, et al.
Veröffentlicht: (2026)
von: Yang, Bangji, et al.
Veröffentlicht: (2026)
MEVG: Multi-event Video Generation with Text-to-Video Models
von: Oh, Gyeongrok, et al.
Veröffentlicht: (2023)
von: Oh, Gyeongrok, et al.
Veröffentlicht: (2023)
Image Generators are Generalist Vision Learners
von: Gabeur, Valentin, et al.
Veröffentlicht: (2026)
von: Gabeur, Valentin, et al.
Veröffentlicht: (2026)
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
von: Oh, Yoonjin, et al.
Veröffentlicht: (2025)
von: Oh, Yoonjin, et al.
Veröffentlicht: (2025)
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
Multi-GRPO: Multi-Group Advantage Estimation for Text-to-Image Generation with Tree-Based Trajectories and Multiple Rewards
von: Lyu, Qiang, et al.
Veröffentlicht: (2025)
von: Lyu, Qiang, et al.
Veröffentlicht: (2025)
NSFW-Classifier Guided Prompt Sanitization for Safe Text-to-Image Generation
von: Xie, Yu, et al.
Veröffentlicht: (2025)
von: Xie, Yu, et al.
Veröffentlicht: (2025)
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
von: Huang, Kaiyi, et al.
Veröffentlicht: (2024)
von: Huang, Kaiyi, et al.
Veröffentlicht: (2024)
Learning Multi-dimensional Human Preference for Text-to-Image Generation
von: Zhang, Sixian, et al.
Veröffentlicht: (2024)
von: Zhang, Sixian, et al.
Veröffentlicht: (2024)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
Conditional Text-to-Image Generation with Reference Guidance
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
von: Oshima, Yuta, et al.
Veröffentlicht: (2025)
von: Oshima, Yuta, et al.
Veröffentlicht: (2025)
Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning
von: Jiang, Kaixun, et al.
Veröffentlicht: (2026)
von: Jiang, Kaixun, et al.
Veröffentlicht: (2026)
Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image Sets
von: Decatur, Dale, et al.
Veröffentlicht: (2025)
von: Decatur, Dale, et al.
Veröffentlicht: (2025)
LayerFusion: Harmonized Multi-Layer Text-to-Image Generation with Generative Priors
von: Dalva, Yusuf, et al.
Veröffentlicht: (2024)
von: Dalva, Yusuf, et al.
Veröffentlicht: (2024)
Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation
von: Lee, Seung Hyun, et al.
Veröffentlicht: (2024)
von: Lee, Seung Hyun, et al.
Veröffentlicht: (2024)
Generative Recall, Dense Reranking: Learning Multi-View Semantic IDs for Efficient Text-to-Video Retrieval
von: Zhao, Zecheng, et al.
Veröffentlicht: (2026)
von: Zhao, Zecheng, et al.
Veröffentlicht: (2026)
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
von: Zeng, Guanning, et al.
Veröffentlicht: (2025)
von: Zeng, Guanning, et al.
Veröffentlicht: (2025)
Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents
von: Song, Yurun, et al.
Veröffentlicht: (2026)
von: Song, Yurun, et al.
Veröffentlicht: (2026)
Escaping Plato's Cave: JAM for Aligning Independently Trained Vision and Language Models
von: Yoon, Lauren Hyoseo, et al.
Veröffentlicht: (2025)
von: Yoon, Lauren Hyoseo, et al.
Veröffentlicht: (2025)
Textual Localization: Decomposing Multi-concept Images for Subject-Driven Text-to-Image Generation
von: Shentu, Junjie, et al.
Veröffentlicht: (2024)
von: Shentu, Junjie, et al.
Veröffentlicht: (2024)
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation
von: Woo, Young Beom, et al.
Veröffentlicht: (2025)
von: Woo, Young Beom, et al.
Veröffentlicht: (2025)
Powerful and Flexible: Personalized Text-to-Image Generation via Reinforcement Learning
von: Wei, Fanyue, et al.
Veröffentlicht: (2024)
von: Wei, Fanyue, et al.
Veröffentlicht: (2024)
Towards Understanding and Quantifying Uncertainty for Text-to-Image Generation
von: Franchi, Gianni, et al.
Veröffentlicht: (2024)
von: Franchi, Gianni, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Aesthetics: Cultural Competence in Text-to-Image Models
von: Kannen, Nithish, et al.
Veröffentlicht: (2024) -
Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing
von: Ma, Shichao, et al.
Veröffentlicht: (2025) -
Alchemist: Turning Public Text-to-Image Data into Generative Gold
von: Startsev, Valerii, et al.
Veröffentlicht: (2025) -
Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits
von: Kalibhat, Neha, et al.
Veröffentlicht: (2026) -
Instilling Multi-round Thinking to Text-guided Image Generation
von: Zeng, Lidong, et al.
Veröffentlicht: (2024)