Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gallici, Matteo, Borde, Haitz Sáez de Ocáriz
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912455568916480
author Gallici, Matteo
Borde, Haitz Sáez de Ocáriz
author_facet Gallici, Matteo
Borde, Haitz Sáez de Ocáriz
contents Fine-tuning pre-trained generative models with Reinforcement Learning (RL) has emerged as an effective approach for aligning outputs more closely with nuanced human preferences. In this paper, we investigate the application of Group Relative Policy Optimization (GRPO) to fine-tune next-scale visual autoregressive (VAR) models. Our empirical results demonstrate that this approach enables alignment to intricate reward signals derived from aesthetic predictors and CLIP embeddings, significantly enhancing image quality and enabling precise control over the generation style. Interestingly, by leveraging CLIP, our method can help VAR models generalize beyond their initial ImageNet distribution: through RL-driven exploration, these models can generate images aligned with prompts referencing image styles that were absent during pre-training. In summary, we show that RL-based fine-tuning is both efficient and effective for VAR models, benefiting particularly from their fast inference speeds, which are advantageous for online sampling, an aspect that poses significant challenges for diffusion-based alternatives.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23331
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization
Gallici, Matteo
Borde, Haitz Sáez de Ocáriz
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Fine-tuning pre-trained generative models with Reinforcement Learning (RL) has emerged as an effective approach for aligning outputs more closely with nuanced human preferences. In this paper, we investigate the application of Group Relative Policy Optimization (GRPO) to fine-tune next-scale visual autoregressive (VAR) models. Our empirical results demonstrate that this approach enables alignment to intricate reward signals derived from aesthetic predictors and CLIP embeddings, significantly enhancing image quality and enabling precise control over the generation style. Interestingly, by leveraging CLIP, our method can help VAR models generalize beyond their initial ImageNet distribution: through RL-driven exploration, these models can generate images aligned with prompts referencing image styles that were absent during pre-training. In summary, we show that RL-based fine-tuning is both efficient and effective for VAR models, benefiting particularly from their fast inference speeds, which are advantageous for online sampling, an aspect that poses significant challenges for diffusion-based alternatives.
title Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.23331