GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Xiong, Tianwei, Liew, Jun Hao, Huang, Zilong, Feng, Jiashi, Liu, Xihui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
di: Xiong, Tianwei, et al.
Pubblicazione: (2026)
di: Xiong, Tianwei, et al.
Pubblicazione: (2026)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
di: Wang, Yuqing, et al.
Pubblicazione: (2024)
di: Wang, Yuqing, et al.
Pubblicazione: (2024)
Parallelized Autoregressive Visual Generation
di: Wang, Yuqing, et al.
Pubblicazione: (2024)
di: Wang, Yuqing, et al.
Pubblicazione: (2024)
ResTok: Learning Hierarchical Residuals in 1D Visual Tokenizers for Autoregressive Image Generation
di: Zhang, Xu, et al.
Pubblicazione: (2026)
di: Zhang, Xu, et al.
Pubblicazione: (2026)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
di: Chen, Cong, et al.
Pubblicazione: (2025)
di: Chen, Cong, et al.
Pubblicazione: (2025)
Empowering Visual Creativity: A Vision-Language Assistant to Image Editing Recommendations
di: Shen, Tiancheng, et al.
Pubblicazione: (2024)
di: Shen, Tiancheng, et al.
Pubblicazione: (2024)
NativeTok: Native Visual Tokenization for Improved Image Generation
di: Wu, Bin, et al.
Pubblicazione: (2026)
di: Wu, Bin, et al.
Pubblicazione: (2026)
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
di: Yue, Yang, et al.
Pubblicazione: (2026)
di: Yue, Yang, et al.
Pubblicazione: (2026)
The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
di: Lei, Weixian, et al.
Pubblicazione: (2025)
di: Lei, Weixian, et al.
Pubblicazione: (2025)
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens
di: Wang, Yuqing, et al.
Pubblicazione: (2026)
di: Wang, Yuqing, et al.
Pubblicazione: (2026)
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
di: Xiong, Tianwei, et al.
Pubblicazione: (2024)
di: Xiong, Tianwei, et al.
Pubblicazione: (2024)
Frequency Autoregressive Image Generation with Continuous Tokens
di: Yu, Hu, et al.
Pubblicazione: (2025)
di: Yu, Hu, et al.
Pubblicazione: (2025)
GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
di: Zhao, Xuan, et al.
Pubblicazione: (2025)
di: Zhao, Xuan, et al.
Pubblicazione: (2025)
DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
di: Zhu, Lianghui, et al.
Pubblicazione: (2024)
di: Zhu, Lianghui, et al.
Pubblicazione: (2024)
Scaling Diffusion Transformers to 16 Billion Parameters
di: Fei, Zhengcong, et al.
Pubblicazione: (2024)
di: Fei, Zhengcong, et al.
Pubblicazione: (2024)
LightningDrag: Lightning Fast and Accurate Drag-based Image Editing Emerging from Videos
di: Shi, Yujun, et al.
Pubblicazione: (2024)
di: Shi, Yujun, et al.
Pubblicazione: (2024)
AlignTok: Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
di: Chen, Bowei, et al.
Pubblicazione: (2025)
di: Chen, Bowei, et al.
Pubblicazione: (2025)
Image Understanding Makes for A Good Tokenizer for Image Generation
di: Wang, Luting, et al.
Pubblicazione: (2024)
di: Wang, Luting, et al.
Pubblicazione: (2024)
ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters
di: Hansen-Estruch, Philippe, et al.
Pubblicazione: (2026)
di: Hansen-Estruch, Philippe, et al.
Pubblicazione: (2026)
ImageFolder: Autoregressive Image Generation with Folded Tokens
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
SuperCLIP: CLIP with Simple Classification Supervision
di: Zhao, Weiheng, et al.
Pubblicazione: (2025)
di: Zhao, Weiheng, et al.
Pubblicazione: (2025)
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
di: Yang, Lihe, et al.
Pubblicazione: (2024)
di: Yang, Lihe, et al.
Pubblicazione: (2024)
TokBench: Evaluating Your Visual Tokenizer before Visual Generation
di: Wu, Junfeng, et al.
Pubblicazione: (2025)
di: Wu, Junfeng, et al.
Pubblicazione: (2025)
DINO-Tok: Adapting DINO for Visual Tokenizers
di: Jia, Mingkai, et al.
Pubblicazione: (2025)
di: Jia, Mingkai, et al.
Pubblicazione: (2025)
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
di: NextStep Team, et al.
Pubblicazione: (2025)
di: NextStep Team, et al.
Pubblicazione: (2025)
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
di: Guo, Yiwei, et al.
Pubblicazione: (2026)
di: Guo, Yiwei, et al.
Pubblicazione: (2026)
RefTok: Reference-Based Tokenization for Video Generation
di: Fan, Xiang, et al.
Pubblicazione: (2025)
di: Fan, Xiang, et al.
Pubblicazione: (2025)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
di: Li, Yan, et al.
Pubblicazione: (2026)
di: Li, Yan, et al.
Pubblicazione: (2026)
Scaling Learned Image Compression Models up to 1 Billion
di: Li, Yuqi, et al.
Pubblicazione: (2025)
di: Li, Yuqi, et al.
Pubblicazione: (2025)
Improving Flexible Image Tokenizers for Autoregressive Image Generation
di: Fu, Zixuan, et al.
Pubblicazione: (2026)
di: Fu, Zixuan, et al.
Pubblicazione: (2026)
MacTok: Robust Continuous Tokenization for Image Generation
di: Zeng, Hengyu, et al.
Pubblicazione: (2026)
di: Zeng, Hengyu, et al.
Pubblicazione: (2026)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
BitDance: Scaling Autoregressive Generative Models with Binary Tokens
di: Ai, Yuang, et al.
Pubblicazione: (2026)
di: Ai, Yuang, et al.
Pubblicazione: (2026)
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
di: Zhang, Luyuan, et al.
Pubblicazione: (2026)
di: Zhang, Luyuan, et al.
Pubblicazione: (2026)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
di: Lin, Huawei, et al.
Pubblicazione: (2025)
di: Lin, Huawei, et al.
Pubblicazione: (2025)
ClusterMark: Towards Robust Watermarking for Autoregressive Image Generators with Visual Token Clustering
di: Lukovnikov, Denis, et al.
Pubblicazione: (2025)
di: Lukovnikov, Denis, et al.
Pubblicazione: (2025)
Classification Done Right for Vision-Language Pre-Training
di: Huang, Zilong, et al.
Pubblicazione: (2024)
di: Huang, Zilong, et al.
Pubblicazione: (2024)
ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
di: Zhang, Kaixin, et al.
Pubblicazione: (2025)
di: Zhang, Kaixin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
di: Xiong, Tianwei, et al.
Pubblicazione: (2026) -
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
di: Wang, Yuqing, et al.
Pubblicazione: (2025) -
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
di: Wang, Yuqing, et al.
Pubblicazione: (2024) -
Parallelized Autoregressive Visual Generation
di: Wang, Yuqing, et al.
Pubblicazione: (2024) -
ResTok: Learning Hierarchical Residuals in 1D Visual Tokenizers for Autoregressive Image Generation
di: Zhang, Xu, et al.
Pubblicazione: (2026)