High-Resolution Image Synthesis via Next-Token Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Dengsheng, Hu, Jie, Yue, Tiezhu, Wei, Xiaoming, Wu, Enhua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Denoising with a Joint-Embedding Predictive Architecture
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
Fine-gained Zero-shot Video Sampling
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
by: Wu, Xiaoping, et al.
Published: (2024)
by: Wu, Xiaoping, et al.
Published: (2024)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
by: Ghasemi, Narges, et al.
Published: (2025)
by: Ghasemi, Narges, et al.
Published: (2025)
ToDo: Token Downsampling for Efficient Generation of High-Resolution Images
by: Smith, Ethan, et al.
Published: (2024)
by: Smith, Ethan, et al.
Published: (2024)
Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers
by: Crowson, Katherine, et al.
Published: (2024)
by: Crowson, Katherine, et al.
Published: (2024)
Segformer++: Efficient Token-Merging Strategies for High-Resolution Semantic Segmentation
by: Kienzle, Daniel, et al.
Published: (2024)
by: Kienzle, Daniel, et al.
Published: (2024)
SkipSR: Faster Super Resolution with Token Skipping
by: Choudhury, Rohan, et al.
Published: (2025)
by: Choudhury, Rohan, et al.
Published: (2025)
Diffusion Autoencoders are Scalable Image Tokenizers
by: Chen, Yinbo, et al.
Published: (2025)
by: Chen, Yinbo, et al.
Published: (2025)
Taming Outlier Tokens in Diffusion Transformers
by: Wu, Xiaoyu, et al.
Published: (2026)
by: Wu, Xiaoyu, et al.
Published: (2026)
Deformable 3D Shape Diffusion Model
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
PoGDiff: Product-of-Gaussians Diffusion Models for Imbalanced Text-to-Image Generation
by: Wang, Ziyan, et al.
Published: (2025)
by: Wang, Ziyan, et al.
Published: (2025)
VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution
by: Zhu, Yixuan, et al.
Published: (2026)
by: Zhu, Yixuan, et al.
Published: (2026)
Improving Diffusion-Based Image Synthesis with Context Prediction
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
ARTA: Adaptive Mixed-Resolution Token Allocation for Efficient Dense Feature Extraction
by: Hagerman, David, et al.
Published: (2026)
by: Hagerman, David, et al.
Published: (2026)
Efficient High-Resolution Image Editing with Hallucination-Aware Loss and Adaptive Tiling
by: Kwon, Young D., et al.
Published: (2025)
by: Kwon, Young D., et al.
Published: (2025)
Next Visual Granularity Generation
by: Wang, Yikai, et al.
Published: (2025)
by: Wang, Yikai, et al.
Published: (2025)
Language-Guided Image Tokenization for Generation
by: Zha, Kaiwen, et al.
Published: (2024)
by: Zha, Kaiwen, et al.
Published: (2024)
Adaptive Length Image Tokenization via Recurrent Allocation
by: Duggal, Shivam, et al.
Published: (2024)
by: Duggal, Shivam, et al.
Published: (2024)
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
by: Gu, Jiatao, et al.
Published: (2025)
by: Gu, Jiatao, et al.
Published: (2025)
Generalizable Geometric Image Caption Synthesis
by: Xin, Yue, et al.
Published: (2025)
by: Xin, Yue, et al.
Published: (2025)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
by: Li, Liupeng, et al.
Published: (2026)
by: Li, Liupeng, et al.
Published: (2026)
Control Your View: High-Resolution Global Semantic Manipulation in Learned Image Compression
by: Liang, Jiaming, et al.
Published: (2026)
by: Liang, Jiaming, et al.
Published: (2026)
Communication-Inspired Tokenization for Structured Image Representations
by: Davtyan, Aram, et al.
Published: (2026)
by: Davtyan, Aram, et al.
Published: (2026)
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
A More Word-like Image Tokenization for MLLMs
by: Lee, Hyun, et al.
Published: (2026)
by: Lee, Hyun, et al.
Published: (2026)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
Single-pass Adaptive Image Tokenization for Minimum Program Search
by: Duggal, Shivam, et al.
Published: (2025)
by: Duggal, Shivam, et al.
Published: (2025)
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
by: Chen, Liang, et al.
Published: (2024)
by: Chen, Liang, et al.
Published: (2024)
Data Overfitting for On-Device Super-Resolution with Dynamic Algorithm and Compiler Co-Design
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
A Study in Dataset Distillation for Image Super-Resolution
by: Dietz, Tobias, et al.
Published: (2025)
by: Dietz, Tobias, et al.
Published: (2025)
Dynamic Attention-Guided Diffusion for Image Super-Resolution
by: Moser, Brian B., et al.
Published: (2023)
by: Moser, Brian B., et al.
Published: (2023)
RefineNet: Enhancing Text-to-Image Conversion with High-Resolution and Detail Accuracy through Hierarchical Transformers and Progressive Refinement
by: Shi, Fan
Published: (2023)
by: Shi, Fan
Published: (2023)
Negative Token Merging: Image-based Adversarial Feature Guidance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
Exploring Token Pruning in Vision State Space Models
by: Zhan, Zheng, et al.
Published: (2024)
by: Zhan, Zheng, et al.
Published: (2024)
The Power of Context: How Multimodality Improves Image Super-Resolution
by: Mei, Kangfu, et al.
Published: (2025)
by: Mei, Kangfu, et al.
Published: (2025)
Hybrid Deep Learning for Hyperspectral Single Image Super-Resolution
by: Muhammad, Usman, et al.
Published: (2025)
by: Muhammad, Usman, et al.
Published: (2025)
Improved Sub-Visible Particle Classification in Flow Imaging Microscopy via Generative AI-Based Image Synthesis
by: Ozbulak, Utku, et al.
Published: (2025)
by: Ozbulak, Utku, et al.
Published: (2025)
Aerial Image Classification in Scarce and Unconstrained Environments via Conformal Prediction
by: Pourkamali-Anaraki, Farhad
Published: (2025)
by: Pourkamali-Anaraki, Farhad
Published: (2025)
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information
by: Wang, Ke, et al.
Published: (2024)
by: Wang, Ke, et al.
Published: (2024)
Similar Items
-
Denoising with a Joint-Embedding Predictive Architecture
by: Chen, Dengsheng, et al.
Published: (2024) -
Fine-gained Zero-shot Video Sampling
by: Chen, Dengsheng, et al.
Published: (2024) -
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
by: Wu, Xiaoping, et al.
Published: (2024) -
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
by: Ghasemi, Narges, et al.
Published: (2025) -
ToDo: Token Downsampling for Efficient Generation of High-Resolution Images
by: Smith, Ethan, et al.
Published: (2024)