Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Kilian, Maciej, Jampani, Varun, Zettlemoyer, Luke |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
di: Zhou, Chunting, et al.
Pubblicazione: (2024)
di: Zhou, Chunting, et al.
Pubblicazione: (2024)
Medical Referring Image Segmentation via Next-Token Mask Prediction
di: Chen, Xinyu, et al.
Pubblicazione: (2025)
di: Chen, Xinyu, et al.
Pubblicazione: (2025)
CAT: Content-Adaptive Image Tokenization
di: Shen, Junhong, et al.
Pubblicazione: (2025)
di: Shen, Junhong, et al.
Pubblicazione: (2025)
HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
di: Hu, Tao, et al.
Pubblicazione: (2026)
di: Hu, Tao, et al.
Pubblicazione: (2026)
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
di: Yang, Shu-wen, et al.
Pubblicazione: (2025)
di: Yang, Shu-wen, et al.
Pubblicazione: (2025)
Object Recognition as Next Token Prediction
di: Yue, Kaiyu, et al.
Pubblicazione: (2023)
di: Yue, Kaiyu, et al.
Pubblicazione: (2023)
High-Resolution Image Synthesis via Next-Token Prediction
di: Chen, Dengsheng, et al.
Pubblicazione: (2024)
di: Chen, Dengsheng, et al.
Pubblicazione: (2024)
NEP: Autoregressive Image Editing via Next Editing Token Prediction
di: Wu, Huimin, et al.
Pubblicazione: (2025)
di: Wu, Huimin, et al.
Pubblicazione: (2025)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
di: Ghasemi, Narges, et al.
Pubblicazione: (2025)
di: Ghasemi, Narges, et al.
Pubblicazione: (2025)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
Unified Dense Prediction of Video Diffusion
di: Yang, Lehan, et al.
Pubblicazione: (2025)
di: Yang, Lehan, et al.
Pubblicazione: (2025)
TokenCompose: Text-to-Image Diffusion with Token-level Supervision
di: Wang, Zirui, et al.
Pubblicazione: (2023)
di: Wang, Zirui, et al.
Pubblicazione: (2023)
Morphing Tokens Draw Strong Masked Image Models
di: Kim, Taekyung, et al.
Pubblicazione: (2023)
di: Kim, Taekyung, et al.
Pubblicazione: (2023)
Token-Space Mask Prediction for Efficient Vision Transformer Segmentation
di: Galagain, Calvin, et al.
Pubblicazione: (2026)
di: Galagain, Calvin, et al.
Pubblicazione: (2026)
EAST: Early Action Prediction Sampling Strategy with Token Masking
di: Sović, Iva, et al.
Pubblicazione: (2026)
di: Sović, Iva, et al.
Pubblicazione: (2026)
Emu3: Next-Token Prediction is All You Need
di: Wang, Xinlong, et al.
Pubblicazione: (2024)
di: Wang, Xinlong, et al.
Pubblicazione: (2024)
Arbitrary Ratio Feature Compression via Next Token Prediction
di: Liu, Yufan, et al.
Pubblicazione: (2026)
di: Liu, Yufan, et al.
Pubblicazione: (2026)
TAPNext: Tracking Any Point (TAP) as Next Token Prediction
di: Zholus, Artem, et al.
Pubblicazione: (2025)
di: Zholus, Artem, et al.
Pubblicazione: (2025)
HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
di: Peng, Xiaogang, et al.
Pubblicazione: (2023)
di: Peng, Xiaogang, et al.
Pubblicazione: (2023)
Humanoid Locomotion as Next Token Prediction
di: Radosavovic, Ilija, et al.
Pubblicazione: (2024)
di: Radosavovic, Ilija, et al.
Pubblicazione: (2024)
Improved Masked Image Generation with Knowledge-Augmented Token Representations
di: Liang, Guotao, et al.
Pubblicazione: (2025)
di: Liang, Guotao, et al.
Pubblicazione: (2025)
SAM-Guided Masked Token Prediction for 3D Scene Understanding
di: Chen, Zhimin, et al.
Pubblicazione: (2024)
di: Chen, Zhimin, et al.
Pubblicazione: (2024)
Cached Adaptive Token Merging: Dynamic Token Reduction and Redundant Computation Elimination in Diffusion Model
di: Saghatchian, Omid, et al.
Pubblicazione: (2025)
di: Saghatchian, Omid, et al.
Pubblicazione: (2025)
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
di: Chen, Weiming, et al.
Pubblicazione: (2026)
di: Chen, Weiming, et al.
Pubblicazione: (2026)
Multimodal Latent Language Modeling with Next-Token Diffusion
di: Sun, Yutao, et al.
Pubblicazione: (2024)
di: Sun, Yutao, et al.
Pubblicazione: (2024)
RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning
di: Xu, Jingqi, et al.
Pubblicazione: (2025)
di: Xu, Jingqi, et al.
Pubblicazione: (2025)
HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Model
di: Nguyen, Hieu T., et al.
Pubblicazione: (2024)
di: Nguyen, Hieu T., et al.
Pubblicazione: (2024)
PanoLlama: Generating Endless and Coherent Panoramas with Next-Token-Prediction LLMs
di: Zhou, Teng, et al.
Pubblicazione: (2024)
di: Zhou, Teng, et al.
Pubblicazione: (2024)
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
di: Chen, Hao, et al.
Pubblicazione: (2025)
di: Chen, Hao, et al.
Pubblicazione: (2025)
AMP: Autoregressive Motion Prediction Revisited with Next Token Prediction for Autonomous Driving
di: Jia, Xiaosong, et al.
Pubblicazione: (2024)
di: Jia, Xiaosong, et al.
Pubblicazione: (2024)
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
di: NextStep Team, et al.
Pubblicazione: (2025)
di: NextStep Team, et al.
Pubblicazione: (2025)
SMooDi: Stylized Motion Diffusion Model
di: Zhong, Lei, et al.
Pubblicazione: (2024)
di: Zhong, Lei, et al.
Pubblicazione: (2024)
SceneStreamer: Continuous Scenario Generation as Next Token Group Prediction
di: Peng, Zhenghao, et al.
Pubblicazione: (2025)
di: Peng, Zhenghao, et al.
Pubblicazione: (2025)
Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions
di: Chen, Lin, et al.
Pubblicazione: (2026)
di: Chen, Lin, et al.
Pubblicazione: (2026)
Rethinking Point Clouds as Sequences: A Causal Next-Token Predictive Learning Framework
di: Yao, Yumeng, et al.
Pubblicazione: (2026)
di: Yao, Yumeng, et al.
Pubblicazione: (2026)
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
di: Zhou, Jensen, et al.
Pubblicazione: (2025)
di: Zhou, Jensen, et al.
Pubblicazione: (2025)
SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video Diffusion
di: Voleti, Vikram, et al.
Pubblicazione: (2024)
di: Voleti, Vikram, et al.
Pubblicazione: (2024)
Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
di: Zhang, Hao, et al.
Pubblicazione: (2025)
di: Zhang, Hao, et al.
Pubblicazione: (2025)
Emerging Property of Masked Token for Effective Pre-training
di: Choi, Hyesong, et al.
Pubblicazione: (2024)
di: Choi, Hyesong, et al.
Pubblicazione: (2024)
ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation
di: Wang, Lingfeng, et al.
Pubblicazione: (2025)
di: Wang, Lingfeng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
di: Zhou, Chunting, et al.
Pubblicazione: (2024) -
Medical Referring Image Segmentation via Next-Token Mask Prediction
di: Chen, Xinyu, et al.
Pubblicazione: (2025) -
CAT: Content-Adaptive Image Tokenization
di: Shen, Junhong, et al.
Pubblicazione: (2025) -
HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
di: Hu, Tao, et al.
Pubblicazione: (2026) -
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
di: Yang, Shu-wen, et al.
Pubblicazione: (2025)