MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jia, Mingkai, Yin, Wei, Hu, Xiaotao, Guo, Jiaxin, Guo, Xiaoyang, Zhang, Qian, Long, Xiao-Xiao, Tan, Ping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Quantize-then-Rectify: Efficient VQ-VAE Training
von: Zhang, Borui, et al.
Veröffentlicht: (2025)
von: Zhang, Borui, et al.
Veröffentlicht: (2025)
Attentive VQ-VAE
von: Hoyos, Angello, et al.
Veröffentlicht: (2023)
von: Hoyos, Angello, et al.
Veröffentlicht: (2023)
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
von: Hu, Xiaotao, et al.
Veröffentlicht: (2024)
von: Hu, Xiaotao, et al.
Veröffentlicht: (2024)
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Exploring VQ-VAE with Prosody Parameters for Speaker Anonymization
von: Leang, Sotheara, et al.
Veröffentlicht: (2024)
von: Leang, Sotheara, et al.
Veröffentlicht: (2024)
Boost 3D Reconstruction using Diffusion-based Monocular Camera Calibration
von: Deng, Junyuan, et al.
Veröffentlicht: (2024)
von: Deng, Junyuan, et al.
Veröffentlicht: (2024)
DINO-Tok: Adapting DINO for Visual Tokenizers
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
ArcVQ-VAE: A Spherical Vector Quantization Framework with ArcCosine Additive Margin
von: Kim, Jaeyung, et al.
Veröffentlicht: (2026)
von: Kim, Jaeyung, et al.
Veröffentlicht: (2026)
2D Gaussians Meet Visual Tokenizer
von: Shi, Yiang, et al.
Veröffentlicht: (2025)
von: Shi, Yiang, et al.
Veröffentlicht: (2025)
PCA-VAE: Differentiable Subspace Quantization without Codebook Collapse
von: Lu, Hao, et al.
Veröffentlicht: (2026)
von: Lu, Hao, et al.
Veröffentlicht: (2026)
Class-Partitioned VQ-VAE and Latent Flow Matching for Point Cloud Scene Generation
von: Edirimuni, Dasith de Silva, et al.
Veröffentlicht: (2026)
von: Edirimuni, Dasith de Silva, et al.
Veröffentlicht: (2026)
MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization
von: Wang, Zhong, et al.
Veröffentlicht: (2026)
von: Wang, Zhong, et al.
Veröffentlicht: (2026)
Flip Distribution Alignment VAE for Multi-Phase MRI Synthesis
von: Kui, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Kui, Xiaoyan, et al.
Veröffentlicht: (2025)
Towards 1000-fold Electron Microscopy Image Compression for Connectomics via VQ-VAE with Transformer Prior
von: Yang, Fuming, et al.
Veröffentlicht: (2025)
von: Yang, Fuming, et al.
Veröffentlicht: (2025)
EO-VAE: Towards A Multi-sensor Tokenizer for Earth Observation Data
von: Lehmann, Nils, et al.
Veröffentlicht: (2026)
von: Lehmann, Nils, et al.
Veröffentlicht: (2026)
D-CNN and VQ-VAE Autoencoders for Compression and Denoising of Industrial X-ray Computed Tomography Images
von: Hejazi, Bardia, et al.
Veröffentlicht: (2025)
von: Hejazi, Bardia, et al.
Veröffentlicht: (2025)
Epsilon-VAE: Denoising as Visual Decoding
von: Zhao, Long, et al.
Veröffentlicht: (2024)
von: Zhao, Long, et al.
Veröffentlicht: (2024)
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
von: Zhao, Sijie, et al.
Veröffentlicht: (2024)
von: Zhao, Sijie, et al.
Veröffentlicht: (2024)
HireVAE: An Online and Adaptive Factor Model Based on Hierarchical and Regime-Switch VAE
von: Wei, Zikai, et al.
Veröffentlicht: (2023)
von: Wei, Zikai, et al.
Veröffentlicht: (2023)
DLFR-VAE: Dynamic Latent Frame Rate VAE for Video Generation
von: Yuan, Zhihang, et al.
Veröffentlicht: (2025)
von: Yuan, Zhihang, et al.
Veröffentlicht: (2025)
CAD-VAE: Leveraging Correlation-Aware Latents for Comprehensive Fair Disentanglement
von: Ma, Chenrui, et al.
Veröffentlicht: (2025)
von: Ma, Chenrui, et al.
Veröffentlicht: (2025)
LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
Asymmetric VAE for One-Step Video Super-Resolution Acceleration
von: Li, Jianze, et al.
Veröffentlicht: (2025)
von: Li, Jianze, et al.
Veröffentlicht: (2025)
Distillation of a tractable model from the VQ-VAE
von: Hadžić, Armin, et al.
Veröffentlicht: (2025)
von: Hadžić, Armin, et al.
Veröffentlicht: (2025)
Precoder Design in Multi-User FDD Systems with VQ-VAE and GNN
von: Allaparapu, Srikar, et al.
Veröffentlicht: (2025)
von: Allaparapu, Srikar, et al.
Veröffentlicht: (2025)
Large Motion Video Autoencoding with Cross-modal Video VAE
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
SepVAE: a contrastive VAE to separate pathological patterns from healthy ones
von: Louiset, Robin, et al.
Veröffentlicht: (2023)
von: Louiset, Robin, et al.
Veröffentlicht: (2023)
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
von: Zhou, Yupeng, et al.
Veröffentlicht: (2025)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2025)
VidTwin: Video VAE with Decoupled Structure and Dynamics
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
OccTENS: 3D Occupancy World Model via Temporal Next-Scale Prediction
von: Jin, Bu, et al.
Veröffentlicht: (2025)
von: Jin, Bu, et al.
Veröffentlicht: (2025)
Feedback Design with VQ-VAE for Robust Precoding in Multi-User FDD Systems
von: Turan, Nurettin, et al.
Veröffentlicht: (2024)
von: Turan, Nurettin, et al.
Veröffentlicht: (2024)
SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation
von: Xing, Ximing, et al.
Veröffentlicht: (2024)
von: Xing, Ximing, et al.
Veröffentlicht: (2024)
Qwen-Image-VAE-2.0 Technical Report
von: Zhang, Zekai, et al.
Veröffentlicht: (2026)
von: Zhang, Zekai, et al.
Veröffentlicht: (2026)
MTC-VAE: Multi-Level Temporal Compression with Content Awareness
von: Dong, Yubo, et al.
Veröffentlicht: (2026)
von: Dong, Yubo, et al.
Veröffentlicht: (2026)
MedVAE: Efficient Automated Interpretation of Medical Images with Large-Scale Generalizable Autoencoders
von: Varma, Maya, et al.
Veröffentlicht: (2025)
von: Varma, Maya, et al.
Veröffentlicht: (2025)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
von: Li, Zongjian, et al.
Veröffentlicht: (2024)
von: Li, Zongjian, et al.
Veröffentlicht: (2024)
VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
von: Bi, Tianci, et al.
Veröffentlicht: (2025)
von: Bi, Tianci, et al.
Veröffentlicht: (2025)
ProbTalk3D: Non-Deterministic Emotion Controllable Speech-Driven 3D Facial Animation Synthesis Using VQ-VAE
von: Wu, Sichun, et al.
Veröffentlicht: (2024)
von: Wu, Sichun, et al.
Veröffentlicht: (2024)
VAE-Var: Variational-Autoencoder-Enhanced Variational Assimilation
von: Xiao, Yi, et al.
Veröffentlicht: (2024)
von: Xiao, Yi, et al.
Veröffentlicht: (2024)
Human Grasp Generation for Rigid and Deformable Objects with Decomposed VQ-VAE
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Quantize-then-Rectify: Efficient VQ-VAE Training
von: Zhang, Borui, et al.
Veröffentlicht: (2025) -
Attentive VQ-VAE
von: Hoyos, Angello, et al.
Veröffentlicht: (2023) -
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
von: Hu, Xiaotao, et al.
Veröffentlicht: (2024) -
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
von: Chen, Hao, et al.
Veröffentlicht: (2024) -
Exploring VQ-VAE with Prosody Parameters for Speaker Anonymization
von: Leang, Sotheara, et al.
Veröffentlicht: (2024)