Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Ping, Zhang, Daoxuan, Wang, Xiangming, Liu, Yungeng, Zeng, Haijin, Chen, Yongyong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:https://arxiv.org/abs/2603.18627
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912973908344832
author Chen, Ping
Zhang, Daoxuan
Wang, Xiangming
Liu, Yungeng
Zeng, Haijin
Chen, Yongyong
author_facet Chen, Ping
Zhang, Daoxuan
Wang, Xiangming
Liu, Yungeng
Zeng, Haijin
Chen, Yongyong
contents Precise Text-to-Image (T2I) generation has achieved great success but is hindered by the limited relational reasoning of static text encoders and the error accumulation in open-loop sampling. Without real-time feedback, initial semantic ambiguities during the Ordinary Differential Equation trajectory inevitably escalate into stochastic deviations from spatial constraints. To bridge this gap, we introduce AFS-Search (Agentic Flow Steering and Parallel Rollout Search), a training-free closed-loop framework built upon FLUX.1-dev. AFS-Search incorporates a training-free closed-loop parallel rollout search and flow steering mechanism, which leverages a Vision-Language Model (VLM) as a semantic critic to diagnose intermediate latents and dynamically steer the velocity field via precise spatial grounding. Complementarily, we formulate T2I generation as a sequential decision-making process, exploring multiple trajectories through lookahead simulations and selecting the optimal path based on VLM-guided rewards. Further, we provide AFS-Search-Pro for higher performance and AFS-Search-Fast for quicker generation. Experimental results show that our AFS-Search-Pro greatly boosts the performance of the original FLUX.1-dev, achieving state-of-the-art results across three different benchmarks. Meanwhile, AFS-Search-Fast also significantly enhances performance while maintaining fast generation speed.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18627
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Agentic Flow Steering and Parallel Rollout Search for Spatially Grounded Text-to-Image Generation
Chen, Ping
Zhang, Daoxuan
Wang, Xiangming
Liu, Yungeng
Zeng, Haijin
Chen, Yongyong
Artificial Intelligence
Precise Text-to-Image (T2I) generation has achieved great success but is hindered by the limited relational reasoning of static text encoders and the error accumulation in open-loop sampling. Without real-time feedback, initial semantic ambiguities during the Ordinary Differential Equation trajectory inevitably escalate into stochastic deviations from spatial constraints. To bridge this gap, we introduce AFS-Search (Agentic Flow Steering and Parallel Rollout Search), a training-free closed-loop framework built upon FLUX.1-dev. AFS-Search incorporates a training-free closed-loop parallel rollout search and flow steering mechanism, which leverages a Vision-Language Model (VLM) as a semantic critic to diagnose intermediate latents and dynamically steer the velocity field via precise spatial grounding. Complementarily, we formulate T2I generation as a sequential decision-making process, exploring multiple trajectories through lookahead simulations and selecting the optimal path based on VLM-guided rewards. Further, we provide AFS-Search-Pro for higher performance and AFS-Search-Fast for quicker generation. Experimental results show that our AFS-Search-Pro greatly boosts the performance of the original FLUX.1-dev, achieving state-of-the-art results across three different benchmarks. Meanwhile, AFS-Search-Fast also significantly enhances performance while maintaining fast generation speed.
title Agentic Flow Steering and Parallel Rollout Search for Spatially Grounded Text-to-Image Generation
topic Artificial Intelligence
url https://arxiv.org/abs/2603.18627