Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Joonhyung, Jang, Hyeongwon, Kim, Joowon, Yang, Eunho |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
von: Kim, Joowon, et al.
Veröffentlicht: (2026)
von: Kim, Joowon, et al.
Veröffentlicht: (2026)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
von: Kim, Sohee, et al.
Veröffentlicht: (2025)
von: Kim, Sohee, et al.
Veröffentlicht: (2025)
PruNeRF: Segment-Centric Dataset Pruning via 3D Spatial Consistency
von: Jung, Yeonsung, et al.
Veröffentlicht: (2024)
von: Jung, Yeonsung, et al.
Veröffentlicht: (2024)
Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing
von: Kim, Yoonjeon, et al.
Veröffentlicht: (2024)
von: Kim, Yoonjeon, et al.
Veröffentlicht: (2024)
LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding
von: Jang, Doohyuk, et al.
Veröffentlicht: (2024)
von: Jang, Doohyuk, et al.
Veröffentlicht: (2024)
Med-PerSAM: One-Shot Visual Prompt Tuning for Personalized Segment Anything Model in Medical Domain
von: Yoon, Hangyul, et al.
Veröffentlicht: (2024)
von: Yoon, Hangyul, et al.
Veröffentlicht: (2024)
Integrating Multimodal Large Language Model Knowledge into Amodal Completion
von: Yun, Heecheol, et al.
Veröffentlicht: (2026)
von: Yun, Heecheol, et al.
Veröffentlicht: (2026)
Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
von: An, Sojung, et al.
Veröffentlicht: (2025)
von: An, Sojung, et al.
Veröffentlicht: (2025)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
von: Tian, Keyu, et al.
Veröffentlicht: (2024)
von: Tian, Keyu, et al.
Veröffentlicht: (2024)
MAGI-1: Autoregressive Video Generation at Scale
von: ai, Sand., et al.
Veröffentlicht: (2025)
von: ai, Sand., et al.
Veröffentlicht: (2025)
BitDance: Scaling Autoregressive Generative Models with Binary Tokens
von: Ai, Yuang, et al.
Veröffentlicht: (2026)
von: Ai, Yuang, et al.
Veröffentlicht: (2026)
I2AM: Interpreting Image-to-Image Latent Diffusion Models via Bi-Attribution Maps
von: Park, Junseo, et al.
Veröffentlicht: (2024)
von: Park, Junseo, et al.
Veröffentlicht: (2024)
Video-T1: Test-Time Scaling for Video Generation
von: Liu, Fangfu, et al.
Veröffentlicht: (2025)
von: Liu, Fangfu, et al.
Veröffentlicht: (2025)
Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
von: Zhang, Zhuoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhuoyang, et al.
Veröffentlicht: (2025)
Direction-Aware Diagonal Autoregressive Image Generation
von: Xu, Yijia, et al.
Veröffentlicht: (2025)
von: Xu, Yijia, et al.
Veröffentlicht: (2025)
Frequency Autoregressive Image Generation with Continuous Tokens
von: Yu, Hu, et al.
Veröffentlicht: (2025)
von: Yu, Hu, et al.
Veröffentlicht: (2025)
Hita: Holistic Tokenizer for Autoregressive Image Generation
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
Visual Accommodation: Rethinking Image Scale as a Learnable Variable for Object Detection
von: Seo, Daeun, et al.
Veröffentlicht: (2024)
von: Seo, Daeun, et al.
Veröffentlicht: (2024)
Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
von: Byun, Ji Young, et al.
Veröffentlicht: (2025)
von: Byun, Ji Young, et al.
Veröffentlicht: (2025)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
von: Xu, Hang, et al.
Veröffentlicht: (2025)
von: Xu, Hang, et al.
Veröffentlicht: (2025)
Scaling Image and Video Generation via Test-Time Evolutionary Search
von: He, Haoran, et al.
Veröffentlicht: (2025)
von: He, Haoran, et al.
Veröffentlicht: (2025)
Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
von: Xiang, Xunzhi, et al.
Veröffentlicht: (2025)
von: Xiang, Xunzhi, et al.
Veröffentlicht: (2025)
BitMark: Watermarking Bitwise Autoregressive Image Generative Models
von: Kerner, Louis, et al.
Veröffentlicht: (2025)
von: Kerner, Louis, et al.
Veröffentlicht: (2025)
Extend3D: Town-Scale 3D Generation
von: Yoon, Seungwoo, et al.
Veröffentlicht: (2026)
von: Yoon, Seungwoo, et al.
Veröffentlicht: (2026)
Training-Free Watermarking for Autoregressive Image Generation
von: Tong, Yu, et al.
Veröffentlicht: (2025)
von: Tong, Yu, et al.
Veröffentlicht: (2025)
TIMING: Temporality-Aware Integrated Gradients for Time Series Explanation
von: Jang, Hyeongwon, et al.
Veröffentlicht: (2025)
von: Jang, Hyeongwon, et al.
Veröffentlicht: (2025)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
von: Ji, Longbin, et al.
Veröffentlicht: (2026)
von: Ji, Longbin, et al.
Veröffentlicht: (2026)
Data-Efficient Unsupervised Interpolation Without Any Intermediate Frame for 4D Medical Images
von: Kim, JungEun, et al.
Veröffentlicht: (2024)
von: Kim, JungEun, et al.
Veröffentlicht: (2024)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
von: Zheng, Zirui, et al.
Veröffentlicht: (2025)
von: Zheng, Zirui, et al.
Veröffentlicht: (2025)
RestoreVAR: Visual Autoregressive Generation for All-in-One Image Restoration
von: Rajagopalan, Sudarshan, et al.
Veröffentlicht: (2025)
von: Rajagopalan, Sudarshan, et al.
Veröffentlicht: (2025)
Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing
von: Kim, Joowon, et al.
Veröffentlicht: (2025)
von: Kim, Joowon, et al.
Veröffentlicht: (2025)
StyleForge: Enhancing Text-to-Image Synthesis for Any Artistic Styles with Dual Binding
von: Park, Junseo, et al.
Veröffentlicht: (2024)
von: Park, Junseo, et al.
Veröffentlicht: (2024)
TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation
von: Shaulov, Ariel, et al.
Veröffentlicht: (2026)
von: Shaulov, Ariel, et al.
Veröffentlicht: (2026)
Automatic Channel Pruning for Multi-Head Attention
von: Lee, Eunho, et al.
Veröffentlicht: (2024)
von: Lee, Eunho, et al.
Veröffentlicht: (2024)
Fast Autoregressive Video Generation with Diagonal Decoding
von: Ye, Yang, et al.
Veröffentlicht: (2025)
von: Ye, Yang, et al.
Veröffentlicht: (2025)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
von: Gu, Zeqi, et al.
Veröffentlicht: (2025)
von: Gu, Zeqi, et al.
Veröffentlicht: (2025)
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
von: Mi, Yapeng, et al.
Veröffentlicht: (2025)
von: Mi, Yapeng, et al.
Veröffentlicht: (2025)
On the Robustness of Watermarking for Autoregressive Image Generation
von: Müller, Andreas, et al.
Veröffentlicht: (2026)
von: Müller, Andreas, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
von: Kim, Joowon, et al.
Veröffentlicht: (2026) -
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
von: Kim, Sohee, et al.
Veröffentlicht: (2025) -
PruNeRF: Segment-Centric Dataset Pruning via 3D Spatial Consistency
von: Jung, Yeonsung, et al.
Veröffentlicht: (2024) -
Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing
von: Kim, Yoonjeon, et al.
Veröffentlicht: (2024) -
LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding
von: Jang, Doohyuk, et al.
Veröffentlicht: (2024)