DiSA: Diffusion Step Annealing in Autoregressive Image Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Qinyu, Singh, Jaskirat, Xu, Ming, Asthana, Akshay, Gould, Stephen, Zheng, Liang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can We Predict Performance of Large Models across Vision-Language Tasks?
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
ARINAR: Bi-Level Autoregressive Feature-by-Feature Generative Models
di: Zhao, Qinyu, et al.
Pubblicazione: (2025)
di: Zhao, Qinyu, et al.
Pubblicazione: (2025)
Towards Optimal Feature-Shaping Methods for Out-of-Distribution Detection
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
Flexible Geometric Guidance for Probabilistic Human Pose Estimation with Diffusion Models
di: Snelgar, Francis, et al.
Pubblicazione: (2026)
di: Snelgar, Francis, et al.
Pubblicazione: (2026)
Gromov Wasserstein Optimal Transport for Semantic Correspondences
di: Snelgar, Francis, et al.
Pubblicazione: (2026)
di: Snelgar, Francis, et al.
Pubblicazione: (2026)
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
di: Wang, Yuchi, et al.
Pubblicazione: (2024)
di: Wang, Yuchi, et al.
Pubblicazione: (2024)
When Diffusion Breaks Constraints: Sequential Autoregressive Generation with RL and MCTS
di: Zhao, Zirui, et al.
Pubblicazione: (2025)
di: Zhao, Zirui, et al.
Pubblicazione: (2025)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
di: Zhao, Qinyu, et al.
Pubblicazione: (2025)
di: Zhao, Qinyu, et al.
Pubblicazione: (2025)
Vec2Face+ for Face Dataset Generation
di: Wu, Haiyu, et al.
Pubblicazione: (2025)
di: Wu, Haiyu, et al.
Pubblicazione: (2025)
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models
di: Xu, Yexing, et al.
Pubblicazione: (2026)
di: Xu, Yexing, et al.
Pubblicazione: (2026)
Vec2Face: Scaling Face Dataset Generation with Loosely Constrained Vectors
di: Wu, Haiyu, et al.
Pubblicazione: (2024)
di: Wu, Haiyu, et al.
Pubblicazione: (2024)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
di: Zhao, Haozhe, et al.
Pubblicazione: (2025)
di: Zhao, Haozhe, et al.
Pubblicazione: (2025)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
di: Kou, Siqi, et al.
Pubblicazione: (2024)
di: Kou, Siqi, et al.
Pubblicazione: (2024)
WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
di: Jiang, Yiwen, et al.
Pubblicazione: (2025)
di: Jiang, Yiwen, et al.
Pubblicazione: (2025)
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
di: Wang, Zihang, et al.
Pubblicazione: (2026)
di: Wang, Zihang, et al.
Pubblicazione: (2026)
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
di: Lyu, Hengye, et al.
Pubblicazione: (2026)
di: Lyu, Hengye, et al.
Pubblicazione: (2026)
Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
di: Huang, Haoyang, et al.
Pubblicazione: (2025)
di: Huang, Haoyang, et al.
Pubblicazione: (2025)
IMPUS: Image Morphing with Perceptually-Uniform Sampling Using Diffusion Models
di: Yang, Zhaoyuan, et al.
Pubblicazione: (2023)
di: Yang, Zhaoyuan, et al.
Pubblicazione: (2023)
REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
di: Leng, Xingjian, et al.
Pubblicazione: (2025)
di: Leng, Xingjian, et al.
Pubblicazione: (2025)
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
di: Zhao, Min, et al.
Pubblicazione: (2026)
di: Zhao, Min, et al.
Pubblicazione: (2026)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
di: Yu, Hao, et al.
Pubblicazione: (2025)
di: Yu, Hao, et al.
Pubblicazione: (2025)
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
di: Chern, Ethan, et al.
Pubblicazione: (2024)
di: Chern, Ethan, et al.
Pubblicazione: (2024)
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
di: Li, Zhuowan, et al.
Pubblicazione: (2024)
di: Li, Zhuowan, et al.
Pubblicazione: (2024)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
ZigMa: A DiT-style Zigzag Mamba Diffusion Model
di: Hu, Vincent Tao, et al.
Pubblicazione: (2024)
di: Hu, Vincent Tao, et al.
Pubblicazione: (2024)
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
di: NextStep Team, et al.
Pubblicazione: (2025)
di: NextStep Team, et al.
Pubblicazione: (2025)
Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression
di: Zhen, Dingcheng, et al.
Pubblicazione: (2025)
di: Zhen, Dingcheng, et al.
Pubblicazione: (2025)
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
di: Zhang, Yu, et al.
Pubblicazione: (2025)
di: Zhang, Yu, et al.
Pubblicazione: (2025)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
di: Chen, Liang, et al.
Pubblicazione: (2025)
di: Chen, Liang, et al.
Pubblicazione: (2025)
Universal Approximation of Visual Autoregressive Transformers
di: Chen, Yifang, et al.
Pubblicazione: (2025)
di: Chen, Yifang, et al.
Pubblicazione: (2025)
Guided Score identity Distillation for Data-Free One-Step Text-to-Image Generation
di: Zhou, Mingyuan, et al.
Pubblicazione: (2024)
di: Zhou, Mingyuan, et al.
Pubblicazione: (2024)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
di: Zhao, Xiangyu, et al.
Pubblicazione: (2023)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2023)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
di: Huang, Jia-Hong, et al.
Pubblicazione: (2024)
di: Huang, Jia-Hong, et al.
Pubblicazione: (2024)
Autoregressive Models in Vision: A Survey
di: Xiong, Jing, et al.
Pubblicazione: (2024)
di: Xiong, Jing, et al.
Pubblicazione: (2024)
Autoregressive Pre-Training on Pixels and Texts
di: Chai, Yekun, et al.
Pubblicazione: (2024)
di: Chai, Yekun, et al.
Pubblicazione: (2024)
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
di: Li, Yizhuo, et al.
Pubblicazione: (2024)
di: Li, Yizhuo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Can We Predict Performance of Large Models across Vision-Language Tasks?
di: Zhao, Qinyu, et al.
Pubblicazione: (2024) -
The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?
di: Zhao, Qinyu, et al.
Pubblicazione: (2024) -
ARINAR: Bi-Level Autoregressive Feature-by-Feature Generative Models
di: Zhao, Qinyu, et al.
Pubblicazione: (2025) -
Towards Optimal Feature-Shaping Methods for Out-of-Distribution Detection
di: Zhao, Qinyu, et al.
Pubblicazione: (2024) -
Flexible Geometric Guidance for Probabilistic Human Pose Estimation with Diffusion Models
di: Snelgar, Francis, et al.
Pubblicazione: (2026)