Scaling Image Tokenizers with Grouped Spherical Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jiangtao, Qin, Zhen, Zhang, Yifan, Hu, Vincent Tao, Ommer, Björn, Briq, Rania, Kesselheim, Stefan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Data Pruning in Generative Diffusion Models
von: Briq, Rania, et al.
Veröffentlicht: (2024)
von: Briq, Rania, et al.
Veröffentlicht: (2024)
[MASK] is All You Need
von: Hu, Vincent Tao, et al.
Veröffentlicht: (2024)
von: Hu, Vincent Tao, et al.
Veröffentlicht: (2024)
TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
von: Krause, Felix, et al.
Veröffentlicht: (2025)
von: Krause, Felix, et al.
Veröffentlicht: (2025)
The Amazing Stability of Flow Matching
von: Briq, Rania, et al.
Veröffentlicht: (2026)
von: Briq, Rania, et al.
Veröffentlicht: (2026)
Exploring and Exploiting Stability in Latent Flow Matching
von: Briq, Rania, et al.
Veröffentlicht: (2026)
von: Briq, Rania, et al.
Veröffentlicht: (2026)
Diffusion Models and Representation Learning: A Survey
von: Fuest, Michael, et al.
Veröffentlicht: (2024)
von: Fuest, Michael, et al.
Veröffentlicht: (2024)
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
von: Matişan, Răzvan-Andrei, et al.
Veröffentlicht: (2025)
von: Matişan, Răzvan-Andrei, et al.
Veröffentlicht: (2025)
Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions
von: Baumann, Stefan Andreas, et al.
Veröffentlicht: (2024)
von: Baumann, Stefan Andreas, et al.
Veröffentlicht: (2024)
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
von: Prestel, Ulrich, et al.
Veröffentlicht: (2026)
von: Prestel, Ulrich, et al.
Veröffentlicht: (2026)
Spherical Leech Quantization for Visual Tokenization and Generation
von: Zhao, Yue, et al.
Veröffentlicht: (2025)
von: Zhao, Yue, et al.
Veröffentlicht: (2025)
ZigMa: A DiT-style Zigzag Mamba Diffusion Model
von: Hu, Vincent Tao, et al.
Veröffentlicht: (2024)
von: Hu, Vincent Tao, et al.
Veröffentlicht: (2024)
Scalable Image Tokenization with Index Backpropagation Quantization
von: Shi, Fengyuan, et al.
Veröffentlicht: (2024)
von: Shi, Fengyuan, et al.
Veröffentlicht: (2024)
MaskFlow: Discrete Flows For Flexible and Efficient Long Video Generation
von: Fuest, Michael, et al.
Veröffentlicht: (2025)
von: Fuest, Michael, et al.
Veröffentlicht: (2025)
Latent Drifting in Diffusion Models for Counterfactual Medical Image Synthesis
von: Yeganeh, Yousef, et al.
Veröffentlicht: (2024)
von: Yeganeh, Yousef, et al.
Veröffentlicht: (2024)
LEPA: Learning Geometric Equivariance in Satellite Remote Sensing Data with a Predictive Architecture
von: Scheurer, Erik, et al.
Veröffentlicht: (2026)
von: Scheurer, Erik, et al.
Veröffentlicht: (2026)
Guiding Token-Sparse Diffusion Models
von: Krause, Felix, et al.
Veröffentlicht: (2026)
von: Krause, Felix, et al.
Veröffentlicht: (2026)
Envisioning the Future, One Step at a Time
von: Baumann, Stefan Andreas, et al.
Veröffentlicht: (2026)
von: Baumann, Stefan Andreas, et al.
Veröffentlicht: (2026)
Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
von: Shao, Run, et al.
Veröffentlicht: (2024)
von: Shao, Run, et al.
Veröffentlicht: (2024)
NoiseDiffusion: Correcting Noise for Image Interpolation with Diffusion Models beyond Spherical Linear Interpolation
von: Zheng, PengFei, et al.
Veröffentlicht: (2024)
von: Zheng, PengFei, et al.
Veröffentlicht: (2024)
MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
FlowCoMotion: Text-to-Motion Generation via Token-Latent Flow Modeling
von: Guan, Dawei, et al.
Veröffentlicht: (2026)
von: Guan, Dawei, et al.
Veröffentlicht: (2026)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
von: Qu, Liao, et al.
Veröffentlicht: (2024)
von: Qu, Liao, et al.
Veröffentlicht: (2024)
Distillation of Diffusion Features for Semantic Correspondence
von: Fundel, Frank, et al.
Veröffentlicht: (2024)
von: Fundel, Frank, et al.
Veröffentlicht: (2024)
VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization
von: Zhang, Yiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Yiwei, et al.
Veröffentlicht: (2024)
Perceptual Group Tokenizer: Building Perception with Iterative Grouping
von: Deng, Zhiwei, et al.
Veröffentlicht: (2023)
von: Deng, Zhiwei, et al.
Veröffentlicht: (2023)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
von: Chen, Cong, et al.
Veröffentlicht: (2025)
von: Chen, Cong, et al.
Veröffentlicht: (2025)
Frequency Autoregressive Image Generation with Continuous Tokens
von: Yu, Hu, et al.
Veröffentlicht: (2025)
von: Yu, Hu, et al.
Veröffentlicht: (2025)
Human-Guided Image Generation for Expanding Small-Scale Training Image Datasets
von: Chen, Changjian, et al.
Veröffentlicht: (2024)
von: Chen, Changjian, et al.
Veröffentlicht: (2024)
Polynomial, trigonometric, and tropical activations
von: Khalfaoui-Hassani, Ismail, et al.
Veröffentlicht: (2025)
von: Khalfaoui-Hassani, Ismail, et al.
Veröffentlicht: (2025)
FASTer: Focal Token Acquiring-and-Scaling Transformer for Long-term 3D Object Detection
von: Dang, Chenxu, et al.
Veröffentlicht: (2025)
von: Dang, Chenxu, et al.
Veröffentlicht: (2025)
Visual Geo-Localization from images
von: Saoud, Rania, et al.
Veröffentlicht: (2024)
von: Saoud, Rania, et al.
Veröffentlicht: (2024)
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition Tokenization
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
BitDance: Scaling Autoregressive Generative Models with Binary Tokens
von: Ai, Yuang, et al.
Veröffentlicht: (2026)
von: Ai, Yuang, et al.
Veröffentlicht: (2026)
Post-Training Quantization for Video Matting
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
von: Ma, Pingchuan, et al.
Veröffentlicht: (2025)
von: Ma, Pingchuan, et al.
Veröffentlicht: (2025)
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
Camyla: Scaling Autonomous Research in Medical Image Segmentation
von: Gao, Yifan, et al.
Veröffentlicht: (2026)
von: Gao, Yifan, et al.
Veröffentlicht: (2026)
Hita: Holistic Tokenizer for Autoregressive Image Generation
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
Vision without Images: End-to-End Computer Vision from Single Compressive Measurements
von: Pan, Fengpu, et al.
Veröffentlicht: (2025)
von: Pan, Fengpu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Data Pruning in Generative Diffusion Models
von: Briq, Rania, et al.
Veröffentlicht: (2024) -
[MASK] is All You Need
von: Hu, Vincent Tao, et al.
Veröffentlicht: (2024) -
TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
von: Krause, Felix, et al.
Veröffentlicht: (2025) -
The Amazing Stability of Flow Matching
von: Briq, Rania, et al.
Veröffentlicht: (2026) -
Exploring and Exploiting Stability in Latent Flow Matching
von: Briq, Rania, et al.
Veröffentlicht: (2026)