Scaling Image Tokenizers with Grouped Spherical Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jiangtao, Qin, Zhen, Zhang, Yifan, Hu, Vincent Tao, Ommer, Björn, Briq, Rania, Kesselheim, Stefan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data Pruning in Generative Diffusion Models
by: Briq, Rania, et al.
Published: (2024)
by: Briq, Rania, et al.
Published: (2024)
[MASK] is All You Need
by: Hu, Vincent Tao, et al.
Published: (2024)
by: Hu, Vincent Tao, et al.
Published: (2024)
TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
by: Krause, Felix, et al.
Published: (2025)
by: Krause, Felix, et al.
Published: (2025)
The Amazing Stability of Flow Matching
by: Briq, Rania, et al.
Published: (2026)
by: Briq, Rania, et al.
Published: (2026)
Exploring and Exploiting Stability in Latent Flow Matching
by: Briq, Rania, et al.
Published: (2026)
by: Briq, Rania, et al.
Published: (2026)
Diffusion Models and Representation Learning: A Survey
by: Fuest, Michael, et al.
Published: (2024)
by: Fuest, Michael, et al.
Published: (2024)
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
by: Matişan, Răzvan-Andrei, et al.
Published: (2025)
by: Matişan, Răzvan-Andrei, et al.
Published: (2025)
Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions
by: Baumann, Stefan Andreas, et al.
Published: (2024)
by: Baumann, Stefan Andreas, et al.
Published: (2024)
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
by: Prestel, Ulrich, et al.
Published: (2026)
by: Prestel, Ulrich, et al.
Published: (2026)
Spherical Leech Quantization for Visual Tokenization and Generation
by: Zhao, Yue, et al.
Published: (2025)
by: Zhao, Yue, et al.
Published: (2025)
ZigMa: A DiT-style Zigzag Mamba Diffusion Model
by: Hu, Vincent Tao, et al.
Published: (2024)
by: Hu, Vincent Tao, et al.
Published: (2024)
Scalable Image Tokenization with Index Backpropagation Quantization
by: Shi, Fengyuan, et al.
Published: (2024)
by: Shi, Fengyuan, et al.
Published: (2024)
MaskFlow: Discrete Flows For Flexible and Efficient Long Video Generation
by: Fuest, Michael, et al.
Published: (2025)
by: Fuest, Michael, et al.
Published: (2025)
Latent Drifting in Diffusion Models for Counterfactual Medical Image Synthesis
by: Yeganeh, Yousef, et al.
Published: (2024)
by: Yeganeh, Yousef, et al.
Published: (2024)
LEPA: Learning Geometric Equivariance in Satellite Remote Sensing Data with a Predictive Architecture
by: Scheurer, Erik, et al.
Published: (2026)
by: Scheurer, Erik, et al.
Published: (2026)
Guiding Token-Sparse Diffusion Models
by: Krause, Felix, et al.
Published: (2026)
by: Krause, Felix, et al.
Published: (2026)
Envisioning the Future, One Step at a Time
by: Baumann, Stefan Andreas, et al.
Published: (2026)
by: Baumann, Stefan Andreas, et al.
Published: (2026)
Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
by: Shao, Run, et al.
Published: (2024)
by: Shao, Run, et al.
Published: (2024)
NoiseDiffusion: Correcting Noise for Image Interpolation with Diffusion Models beyond Spherical Linear Interpolation
by: Zheng, PengFei, et al.
Published: (2024)
by: Zheng, PengFei, et al.
Published: (2024)
MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
FlowCoMotion: Text-to-Motion Generation via Token-Latent Flow Modeling
by: Guan, Dawei, et al.
Published: (2026)
by: Guan, Dawei, et al.
Published: (2026)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
by: Qu, Liao, et al.
Published: (2024)
by: Qu, Liao, et al.
Published: (2024)
Distillation of Diffusion Features for Semantic Correspondence
by: Fundel, Frank, et al.
Published: (2024)
by: Fundel, Frank, et al.
Published: (2024)
VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization
by: Zhang, Yiwei, et al.
Published: (2024)
by: Zhang, Yiwei, et al.
Published: (2024)
Perceptual Group Tokenizer: Building Perception with Iterative Grouping
by: Deng, Zhiwei, et al.
Published: (2023)
by: Deng, Zhiwei, et al.
Published: (2023)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
Frequency Autoregressive Image Generation with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
Human-Guided Image Generation for Expanding Small-Scale Training Image Datasets
by: Chen, Changjian, et al.
Published: (2024)
by: Chen, Changjian, et al.
Published: (2024)
Polynomial, trigonometric, and tropical activations
by: Khalfaoui-Hassani, Ismail, et al.
Published: (2025)
by: Khalfaoui-Hassani, Ismail, et al.
Published: (2025)
FASTer: Focal Token Acquiring-and-Scaling Transformer for Long-term 3D Object Detection
by: Dang, Chenxu, et al.
Published: (2025)
by: Dang, Chenxu, et al.
Published: (2025)
Visual Geo-Localization from images
by: Saoud, Rania, et al.
Published: (2024)
by: Saoud, Rania, et al.
Published: (2024)
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
by: Wang, Xinhao, et al.
Published: (2026)
by: Wang, Xinhao, et al.
Published: (2026)
Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition Tokenization
by: Wang, Cong, et al.
Published: (2025)
by: Wang, Cong, et al.
Published: (2025)
BitDance: Scaling Autoregressive Generative Models with Binary Tokens
by: Ai, Yuang, et al.
Published: (2026)
by: Ai, Yuang, et al.
Published: (2026)
Post-Training Quantization for Video Matting
by: Zhu, Tianrui, et al.
Published: (2025)
by: Zhu, Tianrui, et al.
Published: (2025)
SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
by: Ma, Pingchuan, et al.
Published: (2025)
by: Ma, Pingchuan, et al.
Published: (2025)
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
by: Jia, Chenwei, et al.
Published: (2026)
by: Jia, Chenwei, et al.
Published: (2026)
Camyla: Scaling Autonomous Research in Medical Image Segmentation
by: Gao, Yifan, et al.
Published: (2026)
by: Gao, Yifan, et al.
Published: (2026)
Hita: Holistic Tokenizer for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025)
by: Zheng, Anlin, et al.
Published: (2025)
Vision without Images: End-to-End Computer Vision from Single Compressive Measurements
by: Pan, Fengpu, et al.
Published: (2025)
by: Pan, Fengpu, et al.
Published: (2025)
Similar Items
-
Data Pruning in Generative Diffusion Models
by: Briq, Rania, et al.
Published: (2024) -
[MASK] is All You Need
by: Hu, Vincent Tao, et al.
Published: (2024) -
TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
by: Krause, Felix, et al.
Published: (2025) -
The Amazing Stability of Flow Matching
by: Briq, Rania, et al.
Published: (2026) -
Exploring and Exploiting Stability in Latent Flow Matching
by: Briq, Rania, et al.
Published: (2026)