Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Riise, Erik, Kaya, Mehmet Onurcan, Papadopoulos, Dim P. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Test-Time Scaling for Small Vision-Language Models
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training
by: Lu, Kaixuan, et al.
Published: (2025)
by: Lu, Kaixuan, et al.
Published: (2025)
AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment
by: Lu, Kaixuan, et al.
Published: (2025)
by: Lu, Kaixuan, et al.
Published: (2025)
POEM: Precise Object-level Editing via MLLM control
by: Schouten, Marco, et al.
Published: (2025)
by: Schouten, Marco, et al.
Published: (2025)
DDRM-PR: Fourier Phase Retrieval using Denoising Diffusion Restoration Models
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
I2I-PR: Deep Iterative Refinement for Phase Retrieval using Image-to-Image Diffusion Models
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
prNet: Data-Driven Phase Retrieval via Stochastic Refinement
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
by: Delatolas, Thanos, et al.
Published: (2025)
by: Delatolas, Thanos, et al.
Published: (2025)
Visual Context-Aware Person Fall Detection
by: Nagaj, Aleksander, et al.
Published: (2024)
by: Nagaj, Aleksander, et al.
Published: (2024)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
by: Sun, Peize, et al.
Published: (2024)
by: Sun, Peize, et al.
Published: (2024)
HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement
by: Schouten, Marco, et al.
Published: (2026)
by: Schouten, Marco, et al.
Published: (2026)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
by: Marioriyad, Arash, et al.
Published: (2024)
by: Marioriyad, Arash, et al.
Published: (2024)
WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting
by: Wang, Lezhong, et al.
Published: (2026)
by: Wang, Lezhong, et al.
Published: (2026)
Latent Directions: A Simple Pathway to Bias Mitigation in Generative AI
by: Olmos, Carolina Lopez, et al.
Published: (2024)
by: Olmos, Carolina Lopez, et al.
Published: (2024)
Weak Cube R-CNN: Weakly Supervised 3D Detection using only 2D Bounding Boxes
by: Hansen, Andreas Lau, et al.
Published: (2025)
by: Hansen, Andreas Lau, et al.
Published: (2025)
Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
by: Ma, Nanye, et al.
Published: (2025)
by: Ma, Nanye, et al.
Published: (2025)
Towards High-Quality Image Segmentation: Improving Topology Accuracy by Penalizing Neighbor Pixels
by: Valverde, Juan Miguel, et al.
Published: (2026)
by: Valverde, Juan Miguel, et al.
Published: (2026)
Efficient Conditional Generation on Scale-based Visual Autoregressive Models
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian Conditioning
by: Zhang, Jinhua, et al.
Published: (2025)
by: Zhang, Jinhua, et al.
Published: (2025)
LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
by: Zheng, Hong-Kai, et al.
Published: (2025)
by: Zheng, Hong-Kai, et al.
Published: (2025)
UniMark: Unified Adaptive Multi-bit Watermarking for Autoregressive Image Generators
by: Yilmaz, Yigit, et al.
Published: (2026)
by: Yilmaz, Yigit, et al.
Published: (2026)
Diffusion Beats Autoregressive in Data-Constrained Settings
by: Prabhudesai, Mihir, et al.
Published: (2025)
by: Prabhudesai, Mihir, et al.
Published: (2025)
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
by: Yu, Lijun, et al.
Published: (2023)
by: Yu, Lijun, et al.
Published: (2023)
Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
by: Lei, Haodong, et al.
Published: (2025)
by: Lei, Haodong, et al.
Published: (2025)
Visual Implicit Autoregressive Modeling
by: Jiang, Pengfei, et al.
Published: (2026)
by: Jiang, Pengfei, et al.
Published: (2026)
DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
by: Park, Mingue, et al.
Published: (2025)
by: Park, Mingue, et al.
Published: (2025)
Customize Your Visual Autoregressive Recipe with Set Autoregressive Modeling
by: Liu, Wenze, et al.
Published: (2024)
by: Liu, Wenze, et al.
Published: (2024)
Hybrid Autoregressive-Diffusion Model for Real-Time Sign Language Production
by: Ye, Maoxiao, et al.
Published: (2025)
by: Ye, Maoxiao, et al.
Published: (2025)
Visual Self-Refinement for Autoregressive Models
by: Wang, Jiamian, et al.
Published: (2025)
by: Wang, Jiamian, et al.
Published: (2025)
MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
by: Kristoffersen, Oskar, et al.
Published: (2025)
by: Kristoffersen, Oskar, et al.
Published: (2025)
Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent Modeling
by: Gu, Jiatao, et al.
Published: (2024)
by: Gu, Jiatao, et al.
Published: (2024)
MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
by: Kadambi, Shreya, et al.
Published: (2025)
by: Kadambi, Shreya, et al.
Published: (2025)
ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
by: Zhou, Zhenglin, et al.
Published: (2025)
by: Zhou, Zhenglin, et al.
Published: (2025)
Visual Autoregressive Modeling for Image Super-Resolution
by: Qu, Yunpeng, et al.
Published: (2025)
by: Qu, Yunpeng, et al.
Published: (2025)
Dynamic Mixture-of-Experts for Visual Autoregressive Model
by: Vincenti, Jort, et al.
Published: (2025)
by: Vincenti, Jort, et al.
Published: (2025)
Visual Autoregressive Modelling for Monocular Depth Estimation
by: El-Ghoussani, Amir, et al.
Published: (2025)
by: El-Ghoussani, Amir, et al.
Published: (2025)
Depth Adaptive Efficient Visual Autoregressive Modeling
by: Li, Chunliang, et al.
Published: (2026)
by: Li, Chunliang, et al.
Published: (2026)
CAR: Controllable Autoregressive Modeling for Visual Generation
by: Yao, Ziyu, et al.
Published: (2024)
by: Yao, Ziyu, et al.
Published: (2024)
IE2Video: Adapting Pretrained Diffusion Models for Event-Based Video Reconstruction
by: Torbunov, Dmitrii, et al.
Published: (2025)
by: Torbunov, Dmitrii, et al.
Published: (2025)
D-AR: Diffusion via Autoregressive Models
by: Gao, Ziteng, et al.
Published: (2025)
by: Gao, Ziteng, et al.
Published: (2025)
Similar Items
-
Efficient Test-Time Scaling for Small Vision-Language Models
by: Kaya, Mehmet Onurcan, et al.
Published: (2025) -
Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training
by: Lu, Kaixuan, et al.
Published: (2025) -
AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment
by: Lu, Kaixuan, et al.
Published: (2025) -
POEM: Precise Object-level Editing via MLLM control
by: Schouten, Marco, et al.
Published: (2025) -
DDRM-PR: Fourier Phase Retrieval using Denoising Diffusion Restoration Models
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)