Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gallici, Matteo, Borde, Haitz Sáez de Ocáriz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs
von: Arabpour, Reza, et al.
Veröffentlicht: (2025)
von: Arabpour, Reza, et al.
Veröffentlicht: (2025)
Mathematical Foundations of Geometric Deep Learning
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2025)
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2025)
Metric Learning for Clifford Group Equivariant Neural Networks
von: Ali, Riccardo, et al.
Veröffentlicht: (2024)
von: Ali, Riccardo, et al.
Veröffentlicht: (2024)
k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS
von: De Schouwer, Jonas, et al.
Veröffentlicht: (2026)
von: De Schouwer, Jonas, et al.
Veröffentlicht: (2026)
Keep It Light! Simplifying Image Clustering Via Text-Free Adapters
von: Li, Yicen, et al.
Veröffentlicht: (2025)
von: Li, Yicen, et al.
Veröffentlicht: (2025)
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
von: Zhou, Renping, et al.
Veröffentlicht: (2025)
von: Zhou, Renping, et al.
Veröffentlicht: (2025)
Beyond Parallelism: Synergistic Computational Graph Effects in Multi-Head Attention
von: Borde, Haitz Sáez de Ocáriz
Veröffentlicht: (2025)
von: Borde, Haitz Sáez de Ocáriz
Veröffentlicht: (2025)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
von: Yu, Zhuoran, et al.
Veröffentlicht: (2023)
von: Yu, Zhuoran, et al.
Veröffentlicht: (2023)
Score Distillation via Reparametrized DDIM
von: Lukoianov, Artem, et al.
Veröffentlicht: (2024)
von: Lukoianov, Artem, et al.
Veröffentlicht: (2024)
PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
von: Meng, Ziqiao, et al.
Veröffentlicht: (2025)
von: Meng, Ziqiao, et al.
Veröffentlicht: (2025)
Towards Quantifying Long-Range Interactions in Graph Machine Learning: a Large Graph Dataset and a Measurement
von: Liang, Huidong, et al.
Veröffentlicht: (2025)
von: Liang, Huidong, et al.
Veröffentlicht: (2025)
Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer
von: Ren, Tao, et al.
Veröffentlicht: (2025)
von: Ren, Tao, et al.
Veröffentlicht: (2025)
Fine-Tuning is Fine, if Calibrated
von: Mai, Zheda, et al.
Veröffentlicht: (2024)
von: Mai, Zheda, et al.
Veröffentlicht: (2024)
Selective Mixup Fine-Tuning for Optimizing Non-Decomposable Objectives
von: Ramasubramanian, Shrinivas, et al.
Veröffentlicht: (2024)
von: Ramasubramanian, Shrinivas, et al.
Veröffentlicht: (2024)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
von: Kim, Moo Jin, et al.
Veröffentlicht: (2025)
von: Kim, Moo Jin, et al.
Veröffentlicht: (2025)
Domain-Aware Fine-Tuning of Foundation Models
von: Kaplan, Ugur Ali, et al.
Veröffentlicht: (2024)
von: Kaplan, Ugur Ali, et al.
Veröffentlicht: (2024)
SpectralAR: Spectral Autoregressive Visual Generation
von: Huang, Yuanhui, et al.
Veröffentlicht: (2025)
von: Huang, Yuanhui, et al.
Veröffentlicht: (2025)
Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2025)
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2025)
Next Visual Granularity Generation
von: Wang, Yikai, et al.
Veröffentlicht: (2025)
von: Wang, Yikai, et al.
Veröffentlicht: (2025)
DriveGPT: Scaling Autoregressive Behavior Models for Driving
von: Huang, Xin, et al.
Veröffentlicht: (2024)
von: Huang, Xin, et al.
Veröffentlicht: (2024)
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
von: Tang, Haotian, et al.
Veröffentlicht: (2024)
von: Tang, Haotian, et al.
Veröffentlicht: (2024)
Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual Recognition
von: Mai, Zheda, et al.
Veröffentlicht: (2024)
von: Mai, Zheda, et al.
Veröffentlicht: (2024)
Gradient-based Fine-Tuning through Pre-trained Model Regularization
von: Liu, Xuanbo, et al.
Veröffentlicht: (2025)
von: Liu, Xuanbo, et al.
Veröffentlicht: (2025)
Chart Deep Research in LVLMs via Parallel Relative Policy Optimization
von: Tang, Jiajin, et al.
Veröffentlicht: (2026)
von: Tang, Jiajin, et al.
Veröffentlicht: (2026)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
von: Tian, Keyu, et al.
Veröffentlicht: (2024)
von: Tian, Keyu, et al.
Veröffentlicht: (2024)
Fine-Tuning a Large Vision-Language Model for Artwork's Scoring and Critique
von: Zhang, Zhehan, et al.
Veröffentlicht: (2026)
von: Zhang, Zhehan, et al.
Veröffentlicht: (2026)
RoaD: Rollouts as Demonstrations for Closed-Loop Supervised Fine-Tuning of Autonomous Driving Policies
von: Garcia-Cobo, Guillermo, et al.
Veröffentlicht: (2025)
von: Garcia-Cobo, Guillermo, et al.
Veröffentlicht: (2025)
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
von: Zhang, Jingyi, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2025)
Taming the Entropy Cliff: Variable Codebook Size Quantization for Autoregressive Visual Generation
von: Zheng, Bowen, et al.
Veröffentlicht: (2026)
von: Zheng, Bowen, et al.
Veröffentlicht: (2026)
Visual Autoregressive Transformers Must Use $Ω(n^2 d)$ Memory
von: Cao, Yang, et al.
Veröffentlicht: (2025)
von: Cao, Yang, et al.
Veröffentlicht: (2025)
Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models
von: Tang, Yiwen, et al.
Veröffentlicht: (2023)
von: Tang, Yiwen, et al.
Veröffentlicht: (2023)
Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment
von: Kim, Bryan Sangwoo, et al.
Veröffentlicht: (2025)
von: Kim, Bryan Sangwoo, et al.
Veröffentlicht: (2025)
Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs
von: Hu, Zixuan, et al.
Veröffentlicht: (2024)
von: Hu, Zixuan, et al.
Veröffentlicht: (2024)
FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization
von: Karim, Mohammed Asad, et al.
Veröffentlicht: (2026)
von: Karim, Mohammed Asad, et al.
Veröffentlicht: (2026)
CLEFT: Language-Image Contrastive Learning with Efficient Large Language Model and Prompt Fine-Tuning
von: Du, Yuexi, et al.
Veröffentlicht: (2024)
von: Du, Yuexi, et al.
Veröffentlicht: (2024)
Relational Visual Similarity
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
Universal Approximation of Visual Autoregressive Transformers
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025) -
LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs
von: Arabpour, Reza, et al.
Veröffentlicht: (2025) -
Mathematical Foundations of Geometric Deep Learning
von: Borde, Haitz Sáez de Ocáriz, et al.
Veröffentlicht: (2025) -
Metric Learning for Clifford Group Equivariant Neural Networks
von: Ali, Riccardo, et al.
Veröffentlicht: (2024) -
k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS
von: De Schouwer, Jonas, et al.
Veröffentlicht: (2026)