Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Lei, Wei, Fangyun, Lu, Yanye, Chen, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Text: Frozen Large Language Models in Visual Signal Comprehension
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
SGC-VQGAN: Towards Complex Scene Representation via Semantic Guided Clustering Codebook
von: Ding, Chenjing, et al.
Veröffentlicht: (2024)
von: Ding, Chenjing, et al.
Veröffentlicht: (2024)
Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
von: Chang, Yifan, et al.
Veröffentlicht: (2025)
von: Chang, Yifan, et al.
Veröffentlicht: (2025)
Dual Codebook VQ: Enhanced Image Reconstruction with Reduced Codebook Size
von: Malidarreh, Parisa Boodaghi, et al.
Veröffentlicht: (2025)
von: Malidarreh, Parisa Boodaghi, et al.
Veröffentlicht: (2025)
Physics-Informed Conditional Diffusion for Motion-Robust Retinal Temporal Laser Speckle Contrast Imaging
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
Scribble Hides Class: Promoting Scribble-Based Weakly-Supervised Semantic Segmentation with Its Class Label
von: Zhang, Xinliang, et al.
Veröffentlicht: (2024)
von: Zhang, Xinliang, et al.
Veröffentlicht: (2024)
Training-free Test-time Improvement for Explainable Medical Image Classification
von: He, Hangzhou, et al.
Veröffentlicht: (2025)
von: He, Hangzhou, et al.
Veröffentlicht: (2025)
V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer
von: He, Hangzhou, et al.
Veröffentlicht: (2025)
von: He, Hangzhou, et al.
Veröffentlicht: (2025)
MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance
von: Dong, Yi, et al.
Veröffentlicht: (2025)
von: Dong, Yi, et al.
Veröffentlicht: (2025)
From 100,000+ images to winning the first brain MRI foundation model challenges: Sharing lessons and models
von: Gordaliza, Pedro M., et al.
Veröffentlicht: (2026)
von: Gordaliza, Pedro M., et al.
Veröffentlicht: (2026)
Fast Autoregressive Models for Continuous Latent Generation
von: Hang, Tiankai, et al.
Veröffentlicht: (2025)
von: Hang, Tiankai, et al.
Veröffentlicht: (2025)
Low-Rank Mixture-of-Experts for Continual Medical Image Segmentation
von: Chen, Qian, et al.
Veröffentlicht: (2024)
von: Chen, Qian, et al.
Veröffentlicht: (2024)
Towards Online Continuous Sign Language Recognition and Translation
von: Zuo, Ronglai, et al.
Veröffentlicht: (2024)
von: Zuo, Ronglai, et al.
Veröffentlicht: (2024)
SuperCL: Superpixel Guided Contrastive Learning for Medical Image Segmentation Pre-training
von: Zeng, Shuang, et al.
Veröffentlicht: (2025)
von: Zeng, Shuang, et al.
Veröffentlicht: (2025)
Beyond Stationarity: Rethinking Codebook Collapse in Vector Quantization
von: Lu, Hao, et al.
Veröffentlicht: (2026)
von: Lu, Hao, et al.
Veröffentlicht: (2026)
DicFace: Dirichlet-Constrained Variational Codebook Learning for Temporally Coherent Video Face Restoration
von: Chen, Yan, et al.
Veröffentlicht: (2025)
von: Chen, Yan, et al.
Veröffentlicht: (2025)
MCR-VQGAN: A Scalable and Cost-Effective Tau PET Synthesis Approach for Alzheimer's Disease Imaging
von: Kim, Jin Young, et al.
Veröffentlicht: (2025)
von: Kim, Jin Young, et al.
Veröffentlicht: (2025)
Bridging Degradation Discrimination and Generation for Universal Image Restoration
von: Hu, JiaKui, et al.
Veröffentlicht: (2026)
von: Hu, JiaKui, et al.
Veröffentlicht: (2026)
Universal Image Restoration Pre-training via Degradation Classification
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
Universal Image Restoration Pre-training via Masked Degradation Classification
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration
von: Yao, Zhengjian, et al.
Veröffentlicht: (2026)
von: Yao, Zhengjian, et al.
Veröffentlicht: (2026)
Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models
von: He, Hangzhou, et al.
Veröffentlicht: (2025)
von: He, Hangzhou, et al.
Veröffentlicht: (2025)
Exploiting Inherent Class Label: Towards Robust Scribble Supervised Semantic Segmentation
von: Zhang, Xinliang, et al.
Veröffentlicht: (2025)
von: Zhang, Xinliang, et al.
Veröffentlicht: (2025)
Inter- and Intra-image Refinement for Few Shot Segmentation
von: Fu, Ourui, et al.
Veröffentlicht: (2025)
von: Fu, Ourui, et al.
Veröffentlicht: (2025)
RADA: Region-Aware Dual-encoder Auxiliary learning for Barely-supervised Medical Image Segmentation
von: Zeng, Shuang, et al.
Veröffentlicht: (2026)
von: Zeng, Shuang, et al.
Veröffentlicht: (2026)
GLARE: Low Light Image Enhancement via Generative Latent Feature based Codebook Retrieval
von: Zhou, Han, et al.
Veröffentlicht: (2024)
von: Zhou, Han, et al.
Veröffentlicht: (2024)
A Hong Kong Sign Language Corpus Collected from Sign-interpreted TV News
von: Niu, Zhe, et al.
Veröffentlicht: (2024)
von: Niu, Zhe, et al.
Veröffentlicht: (2024)
From Virtual Games to Real-World Play
von: Sun, Wenqiang, et al.
Veröffentlicht: (2025)
von: Sun, Wenqiang, et al.
Veröffentlicht: (2025)
HybridFlow: Infusing Continuity into Masked Codebook for Extreme Low-Bitrate Image Compression
von: Lu, Lei, et al.
Veröffentlicht: (2024)
von: Lu, Lei, et al.
Veröffentlicht: (2024)
A Simple Baseline for Spoken Language to Sign Language Translation with 3D Avatars
von: Zuo, Ronglai, et al.
Veröffentlicht: (2024)
von: Zuo, Ronglai, et al.
Veröffentlicht: (2024)
Animate Any Character in Any World
von: Wang, Yitong, et al.
Veröffentlicht: (2025)
von: Wang, Yitong, et al.
Veröffentlicht: (2025)
Auto-Regressively Generating Multi-View Consistent Images
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
Enhancing Image Restoration Transformer via Adaptive Translation Equivariance
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation
von: Zhao, Shuling, et al.
Veröffentlicht: (2024)
von: Zhao, Shuling, et al.
Veröffentlicht: (2024)
MCGA: Mixture of Codebooks Hyperspectral Reconstruction via Grayscale-Aware Attention
von: Yang, Zhanjiang, et al.
Veröffentlicht: (2025)
von: Yang, Zhanjiang, et al.
Veröffentlicht: (2025)
Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long Text
von: Liang, Guotao, et al.
Veröffentlicht: (2025)
von: Liang, Guotao, et al.
Veröffentlicht: (2025)
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
von: Yue, Yang, et al.
Veröffentlicht: (2026)
von: Yue, Yang, et al.
Veröffentlicht: (2026)
PCA-VAE: Differentiable Subspace Quantization without Codebook Collapse
von: Lu, Hao, et al.
Veröffentlicht: (2026)
von: Lu, Hao, et al.
Veröffentlicht: (2026)
AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs
von: Zhang, Xinliang, et al.
Veröffentlicht: (2025)
von: Zhang, Xinliang, et al.
Veröffentlicht: (2025)
Expressive and Generalizable Low-rank Adaptation for Large Models via Slow Cascaded Learning
von: Li, Siwei, et al.
Veröffentlicht: (2024)
von: Li, Siwei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Text: Frozen Large Language Models in Visual Signal Comprehension
von: Zhu, Lei, et al.
Veröffentlicht: (2024) -
SGC-VQGAN: Towards Complex Scene Representation via Semantic Guided Clustering Codebook
von: Ding, Chenjing, et al.
Veröffentlicht: (2024) -
Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
von: Chang, Yifan, et al.
Veröffentlicht: (2025) -
Dual Codebook VQ: Enhanced Image Reconstruction with Reduced Codebook Size
von: Malidarreh, Parisa Boodaghi, et al.
Veröffentlicht: (2025) -
Physics-Informed Conditional Diffusion for Motion-Robust Retinal Temporal Laser Speckle Contrast Imaging
von: Chen, Qian, et al.
Veröffentlicht: (2026)