An Image is Worth 32 Tokens for Reconstruction and Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Qihang, Weber, Mark, Deng, Xueqing, Shen, Xiaohui, Cremers, Daniel, Chen, Liang-Chieh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MaskBit: Embedding-free Image Generation via Bit Tokens
di: Weber, Mark, et al.
Pubblicazione: (2024)
di: Weber, Mark, et al.
Pubblicazione: (2024)
Randomized Autoregressive Visual Generation
di: Yu, Qihang, et al.
Pubblicazione: (2024)
di: Yu, Qihang, et al.
Pubblicazione: (2024)
COCONut: Modernizing COCO Segmentation
di: Deng, Xueqing, et al.
Pubblicazione: (2024)
di: Deng, Xueqing, et al.
Pubblicazione: (2024)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
di: Kerssies, Tommie, et al.
Pubblicazione: (2026)
di: Kerssies, Tommie, et al.
Pubblicazione: (2026)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
di: Kim, Dongwon, et al.
Pubblicazione: (2025)
di: Kim, Dongwon, et al.
Pubblicazione: (2025)
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
di: He, Ju, et al.
Pubblicazione: (2023)
di: He, Ju, et al.
Pubblicazione: (2023)
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
di: Deng, Xueqing, et al.
Pubblicazione: (2025)
di: Deng, Xueqing, et al.
Pubblicazione: (2025)
Frequency-Aware Flow Matching for High-Quality Image Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2026)
di: Ren, Sucheng, et al.
Pubblicazione: (2026)
FlowTok: Flowing Seamlessly Across Text and Image Tokens
di: He, Ju, et al.
Pubblicazione: (2025)
di: He, Ju, et al.
Pubblicazione: (2025)
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization
di: Liu, Qihao, et al.
Pubblicazione: (2024)
di: Liu, Qihao, et al.
Pubblicazione: (2024)
Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting
di: Shin, Inkyu, et al.
Pubblicazione: (2024)
di: Shin, Inkyu, et al.
Pubblicazione: (2024)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
di: Chen, Jieneng, et al.
Pubblicazione: (2024)
di: Chen, Jieneng, et al.
Pubblicazione: (2024)
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
di: Athar, Ali, et al.
Pubblicazione: (2024)
di: Athar, Ali, et al.
Pubblicazione: (2024)
Autoregressive Image Generation with Masked Bit Modeling
di: Yu, Qihang, et al.
Pubblicazione: (2026)
di: Yu, Qihang, et al.
Pubblicazione: (2026)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
di: Wang, Feng, et al.
Pubblicazione: (2025)
di: Wang, Feng, et al.
Pubblicazione: (2025)
1.58-bit FLUX
di: Yang, Chenglin, et al.
Pubblicazione: (2024)
di: Yang, Chenglin, et al.
Pubblicazione: (2024)
Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images
di: Wulff, Philipp, et al.
Pubblicazione: (2025)
di: Wulff, Philipp, et al.
Pubblicazione: (2025)
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
di: Hu, Lianyu, et al.
Pubblicazione: (2024)
di: Hu, Lianyu, et al.
Pubblicazione: (2024)
Images are Worth Variable Length of Representations
di: Mao, Lingjun, et al.
Pubblicazione: (2025)
di: Mao, Lingjun, et al.
Pubblicazione: (2025)
Semantic One-Dimensional Tokenizer for Image Reconstruction and Generation
di: Qu, Yunpeng, et al.
Pubblicazione: (2026)
di: Qu, Yunpeng, et al.
Pubblicazione: (2026)
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
di: Liu, Qihao, et al.
Pubblicazione: (2025)
di: Liu, Qihao, et al.
Pubblicazione: (2025)
Power Variable Projection for Initialization-Free Large-Scale Bundle Adjustment
di: Weber, Simon, et al.
Pubblicazione: (2024)
di: Weber, Simon, et al.
Pubblicazione: (2024)
Flattening the Parent Bias: Hierarchical Semantic Segmentation in the Poincaré Ball
di: Weber, Simon, et al.
Pubblicazione: (2024)
di: Weber, Simon, et al.
Pubblicazione: (2024)
Finsler-Laplace-Beltrami Operators with Application to Shape Analysis
di: Weber, Simon, et al.
Pubblicazione: (2024)
di: Weber, Simon, et al.
Pubblicazione: (2024)
A Creative Agent is Worth a 64-Token Template
di: Shi, Ruixiao, et al.
Pubblicazione: (2026)
di: Shi, Ruixiao, et al.
Pubblicazione: (2026)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
di: Chen, Cong, et al.
Pubblicazione: (2025)
di: Chen, Cong, et al.
Pubblicazione: (2025)
Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens
di: Xie, Qingsong, et al.
Pubblicazione: (2025)
di: Xie, Qingsong, et al.
Pubblicazione: (2025)
Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
di: Qian, Shenhan, et al.
Pubblicazione: (2026)
di: Qian, Shenhan, et al.
Pubblicazione: (2026)
GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
di: Zhao, Xuan, et al.
Pubblicazione: (2025)
di: Zhao, Xuan, et al.
Pubblicazione: (2025)
NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
di: Chen, Weirong, et al.
Pubblicazione: (2026)
di: Chen, Weirong, et al.
Pubblicazione: (2026)
A Pixel Is Worth More Than One 3D Gaussians in Single-View 3D Reconstruction
di: Shen, Jianghao, et al.
Pubblicazione: (2024)
di: Shen, Jianghao, et al.
Pubblicazione: (2024)
TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection
di: Zhou, Qihang, et al.
Pubblicazione: (2025)
di: Zhou, Qihang, et al.
Pubblicazione: (2025)
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
di: Fan, Qihang, et al.
Pubblicazione: (2024)
di: Fan, Qihang, et al.
Pubblicazione: (2024)
Generative Shape Reconstruction with Geometry-Guided Langevin Dynamics
di: Härenstam-Nielsen, Linus, et al.
Pubblicazione: (2026)
di: Härenstam-Nielsen, Linus, et al.
Pubblicazione: (2026)
A Label is Worth a Thousand Images in Dataset Distillation
di: Qin, Tian, et al.
Pubblicazione: (2024)
di: Qin, Tian, et al.
Pubblicazione: (2024)
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
Back on Track: Bundle Adjustment for Dynamic Scene Reconstruction
di: Chen, Weirong, et al.
Pubblicazione: (2025)
di: Chen, Weirong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MaskBit: Embedding-free Image Generation via Bit Tokens
di: Weber, Mark, et al.
Pubblicazione: (2024) -
Randomized Autoregressive Visual Generation
di: Yu, Qihang, et al.
Pubblicazione: (2024) -
COCONut: Modernizing COCO Segmentation
di: Deng, Xueqing, et al.
Pubblicazione: (2024) -
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
di: Kerssies, Tommie, et al.
Pubblicazione: (2026) -
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2025)