Rethinking Test Time Scaling for Flow-Matching Generative Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yu, Qingtao, Song, Changlin, Sun, Minghao, Yu, Zhengyang, Verma, Vinay Kumar, Roy, Soumya, Negi, Sumit, Li, Hongdong, Campbell, Dylan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918398695309312
author Yu, Qingtao
Song, Changlin
Sun, Minghao
Yu, Zhengyang
Verma, Vinay Kumar
Roy, Soumya
Negi, Sumit
Li, Hongdong
Campbell, Dylan
author_facet Yu, Qingtao
Song, Changlin
Sun, Minghao
Yu, Zhengyang
Verma, Vinay Kumar
Roy, Soumya
Negi, Sumit
Li, Hongdong
Campbell, Dylan
contents The performance of text-to-image diffusion models may be improved at test-time by scaling computation to search for a generated image that maximizes a given reward function. While existing trajectory level exploration methods improve the effectiveness of test-time scaling for standard diffusion models, they are largely incompatible with modern flow matching models, which use deterministic sampling. This imposes significant computational overhead on local trajectory search, making the trade-offs less favorable compared to global search. However, global search strategies like trajectory pruning face two critical challenges: the sharp, low-diversity distributions characteristic of scaled flow models that restrict the candidate search space, and the bias of reward models in the early denoising process. To overcome these limitations, we propose Repel, a token-level mechanism that encourages sample diversity, and NARF, a noise-aware reward fine-tuning strategy to obtain more accurate reward ranking at early denoising stages. Together, these promote more effective test-time scaling resource allocation. Overall, we name our pipeline as \textbf{DOG-Trim}: \textbf{D}iversity enhanced \textbf{O}rder aligned \textbf{G}lobal flow Trimming. The experiments demonstrate that, under the same compute cost, our approach achieves around twice the performance improvement relative to the scaling-free baseline compared to the best existing method. Github: https://github.com/TerrysLearning/DOGTrimTTS.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22242
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Test Time Scaling for Flow-Matching Generative Models
Yu, Qingtao
Song, Changlin
Sun, Minghao
Yu, Zhengyang
Verma, Vinay Kumar
Roy, Soumya
Negi, Sumit
Li, Hongdong
Campbell, Dylan
Computer Vision and Pattern Recognition
The performance of text-to-image diffusion models may be improved at test-time by scaling computation to search for a generated image that maximizes a given reward function. While existing trajectory level exploration methods improve the effectiveness of test-time scaling for standard diffusion models, they are largely incompatible with modern flow matching models, which use deterministic sampling. This imposes significant computational overhead on local trajectory search, making the trade-offs less favorable compared to global search. However, global search strategies like trajectory pruning face two critical challenges: the sharp, low-diversity distributions characteristic of scaled flow models that restrict the candidate search space, and the bias of reward models in the early denoising process. To overcome these limitations, we propose Repel, a token-level mechanism that encourages sample diversity, and NARF, a noise-aware reward fine-tuning strategy to obtain more accurate reward ranking at early denoising stages. Together, these promote more effective test-time scaling resource allocation. Overall, we name our pipeline as \textbf{DOG-Trim}: \textbf{D}iversity enhanced \textbf{O}rder aligned \textbf{G}lobal flow Trimming. The experiments demonstrate that, under the same compute cost, our approach achieves around twice the performance improvement relative to the scaling-free baseline compared to the best existing method. Github: https://github.com/TerrysLearning/DOGTrimTTS.
title Rethinking Test Time Scaling for Flow-Matching Generative Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.22242