DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Mingue, Phunyaphibarn, Prin, Lee, Phillip Y., Sung, Minhyuk
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911288232247296
author Park, Mingue
Phunyaphibarn, Prin
Lee, Phillip Y.
Sung, Minhyuk
author_facet Park, Mingue
Phunyaphibarn, Prin
Lee, Phillip Y.
Sung, Minhyuk
contents We introduce DiverseVAR, a framework that enhances the diversity of text-conditioned visual autoregressive models (VAR) at test time without requiring retraining, fine-tuning, or substantial computational overhead. While VAR models have recently emerged as strong competitors to diffusion and flow models for image generation, they suffer from a critical limitation in diversity, often producing nearly identical images even for simple prompts. This issue has largely gone unnoticed amid the predominant focus on image quality. We address this limitation at test time in two stages. First, inspired by diversity enhancement techniques in diffusion models, we propose injecting noise into the text embedding. This introduces a trade-off between diversity and image quality: as diversity increases, the image quality sharply declines. To preserve quality, we propose scale-travel: a novel latent refinement technique inspired by time-travel strategies in diffusion models. Specifically, we use a multi-scale autoencoder to extract coarse-scale tokens that enable us to resume generation at intermediate stages. Extensive experiments show that combining text-embedding noise injection with our scale-travel refinement significantly enhances diversity while minimizing image-quality degradation, achieving a new Pareto frontier in the diversity-quality trade-off.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21415
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
Park, Mingue
Phunyaphibarn, Prin
Lee, Phillip Y.
Sung, Minhyuk
Computer Vision and Pattern Recognition
We introduce DiverseVAR, a framework that enhances the diversity of text-conditioned visual autoregressive models (VAR) at test time without requiring retraining, fine-tuning, or substantial computational overhead. While VAR models have recently emerged as strong competitors to diffusion and flow models for image generation, they suffer from a critical limitation in diversity, often producing nearly identical images even for simple prompts. This issue has largely gone unnoticed amid the predominant focus on image quality. We address this limitation at test time in two stages. First, inspired by diversity enhancement techniques in diffusion models, we propose injecting noise into the text embedding. This introduces a trade-off between diversity and image quality: as diversity increases, the image quality sharply declines. To preserve quality, we propose scale-travel: a novel latent refinement technique inspired by time-travel strategies in diffusion models. Specifically, we use a multi-scale autoencoder to extract coarse-scale tokens that enable us to resume generation at intermediate stages. Extensive experiments show that combining text-embedding noise injection with our scale-travel refinement significantly enhances diversity while minimizing image-quality degradation, achieving a new Pareto frontier in the diversity-quality trade-off.
title DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.21415