Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Xingchen, Zhou, Han, Sun, Ruoxi, Nakhost, Hootan, Jiang, Ke, Sinha, Rajarishi, Arık, Sercan Ö. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and Generation
by: Wan, Xingchen, et al.
Published: (2025)
by: Wan, Xingchen, et al.
Published: (2025)
Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt Optimization
by: Wan, Xingchen, et al.
Published: (2024)
by: Wan, Xingchen, et al.
Published: (2024)
VISTA: A Test-Time Self-Improving Video Generation Agent
by: Long, Do Xuan, et al.
Published: (2025)
by: Long, Do Xuan, et al.
Published: (2025)
SQL-PaLM: Improved Large Language Model Adaptation for Text-to-SQL (extended)
by: Sun, Ruoxi, et al.
Published: (2023)
by: Sun, Ruoxi, et al.
Published: (2023)
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies
by: Zhou, Han, et al.
Published: (2025)
by: Zhou, Han, et al.
Published: (2025)
Effective Large Language Model Adaptation for Improved Grounding and Citation Generation
by: Ye, Xi, et al.
Published: (2023)
by: Ye, Xi, et al.
Published: (2023)
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
by: Chen, Maximillian, et al.
Published: (2024)
by: Chen, Maximillian, et al.
Published: (2024)
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
by: Chen, Maximillian, et al.
Published: (2024)
by: Chen, Maximillian, et al.
Published: (2024)
Reasoning-SQL: Reinforcement Learning with SQL Tailored Partial Rewards for Reasoning-Enhanced Text-to-SQL
by: Pourreza, Mohammadreza, et al.
Published: (2025)
by: Pourreza, Mohammadreza, et al.
Published: (2025)
Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments
by: Su, Hongjin, et al.
Published: (2025)
by: Su, Hongjin, et al.
Published: (2025)
SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling
by: Chen, Jiefeng, et al.
Published: (2025)
by: Chen, Jiefeng, et al.
Published: (2025)
SQL-GEN: Bridging the Dialect Gap for Text-to-SQL Via Synthetic Data And Model Merging
by: Pourreza, Mohammadreza, et al.
Published: (2024)
by: Pourreza, Mohammadreza, et al.
Published: (2024)
Chain of Agents: Large Language Models Collaborating on Long-Context Tasks
by: Zhang, Yusen, et al.
Published: (2024)
by: Zhang, Yusen, et al.
Published: (2024)
CROME: Cross-Modal Adapters for Efficient Multimodal LLM
by: Ebrahimi, Sayna, et al.
Published: (2024)
by: Ebrahimi, Sayna, et al.
Published: (2024)
Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment
by: Sarkar, Pritam, et al.
Published: (2024)
by: Sarkar, Pritam, et al.
Published: (2024)
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation
by: Chen, Sixiang, et al.
Published: (2026)
by: Chen, Sixiang, et al.
Published: (2026)
FLAIRR-TS -- Forecasting LLM-Agents with Iterative Refinement and Retrieval for Time Series
by: Jalori, Gunjan, et al.
Published: (2025)
by: Jalori, Gunjan, et al.
Published: (2025)
Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding Dimensions
by: Yoon, Jinsung, et al.
Published: (2024)
by: Yoon, Jinsung, et al.
Published: (2024)
COSTAR: Improved Temporal Counterfactual Estimation with Self-Supervised Learning
by: Meng, Chuizheng, et al.
Published: (2023)
by: Meng, Chuizheng, et al.
Published: (2023)
Thinking with Images via Self-Calling Agent
by: Yang, Wenxi, et al.
Published: (2025)
by: Yang, Wenxi, et al.
Published: (2025)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning
by: Jiang, Kaixun, et al.
Published: (2026)
by: Jiang, Kaixun, et al.
Published: (2026)
SSVIF: Self-Supervised Segmentation-Oriented Visible and Infrared Image Fusion
by: Zhao, Zixian, et al.
Published: (2025)
by: Zhao, Zixian, et al.
Published: (2025)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
by: Kim, Seungwook, et al.
Published: (2026)
by: Kim, Seungwook, et al.
Published: (2026)
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
by: Yang, Hongji, et al.
Published: (2025)
by: Yang, Hongji, et al.
Published: (2025)
VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation
by: Jiang, Kaiyuan, et al.
Published: (2025)
by: Jiang, Kaiyuan, et al.
Published: (2025)
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
by: Jin, Bowen, et al.
Published: (2025)
by: Jin, Bowen, et al.
Published: (2025)
Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA
by: Ke, Yan, et al.
Published: (2025)
by: Ke, Yan, et al.
Published: (2025)
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
by: Oh, Yoonjin, et al.
Published: (2025)
by: Oh, Yoonjin, et al.
Published: (2025)
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
by: Guo, Yuxiang, et al.
Published: (2025)
by: Guo, Yuxiang, et al.
Published: (2025)
Text4Seg: Reimagining Image Segmentation as Text Generation
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
by: Lan, Mengcheng, et al.
Published: (2025)
by: Lan, Mengcheng, et al.
Published: (2025)
Boosting All-in-One Image Restoration via Self-Improved Privilege Learning
by: Wu, Gang, et al.
Published: (2025)
by: Wu, Gang, et al.
Published: (2025)
MultiRef: Controllable Image Generation with Multiple Visual References
by: Chen, Ruoxi, et al.
Published: (2025)
by: Chen, Ruoxi, et al.
Published: (2025)
Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings
by: Zarei, Arman, et al.
Published: (2024)
by: Zarei, Arman, et al.
Published: (2024)
Large Language Models Can Automatically Engineer Features for Few-Shot Tabular Learning
by: Han, Sungwon, et al.
Published: (2024)
by: Han, Sungwon, et al.
Published: (2024)
Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG
by: Jin, Bowen, et al.
Published: (2024)
by: Jin, Bowen, et al.
Published: (2024)
Detect-and-Guide: Self-regulation of Diffusion Models for Safe Text-to-Image Generation via Guideline Token Optimization
by: Li, Feifei, et al.
Published: (2025)
by: Li, Feifei, et al.
Published: (2025)
Similar Items
-
From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and Generation
by: Wan, Xingchen, et al.
Published: (2025) -
Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt Optimization
by: Wan, Xingchen, et al.
Published: (2024) -
VISTA: A Test-Time Self-Improving Video Generation Agent
by: Long, Do Xuan, et al.
Published: (2025) -
SQL-PaLM: Improved Large Language Model Adaptation for Text-to-SQL (extended)
by: Sun, Ruoxi, et al.
Published: (2023) -
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
by: Wang, Fei, et al.
Published: (2025)