Image and Video Tokenization with Binary Spherical Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yue, Xiong, Yuanjun, Krähenbühl, Philipp |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GANCompress: GAN-Enhanced Neural Image Compression with Binary Spherical Quantization
by: Sivakoti, Karthik
Published: (2025)
by: Sivakoti, Karthik
Published: (2025)
Bagged Deep Image Prior for Recovering Images in the Presence of Speckle Noise
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
High Perceptual Quality Wireless Image Delivery with Denoising Diffusion Models
by: Yilmaz, Selim F., et al.
Published: (2023)
by: Yilmaz, Selim F., et al.
Published: (2023)
Quantifying Knowledge Distillation Using Partial Information Decomposition
by: Dissanayake, Pasan, et al.
Published: (2024)
by: Dissanayake, Pasan, et al.
Published: (2024)
Fast and Accurate Cooperative Radio Map Estimation Enabled by GAN
by: Zhang, Zezhong, et al.
Published: (2024)
by: Zhang, Zezhong, et al.
Published: (2024)
Laplacian-guided Entropy Model in Neural Codec with Blur-dissipated Synthesis
by: Khoshkhahtinat, Atefeh, et al.
Published: (2024)
by: Khoshkhahtinat, Atefeh, et al.
Published: (2024)
Resi-VidTok: An Efficient and Decomposed Progressive Tokenization Framework for Ultra-Low-Rate and Lightweight Video Transmission
by: Liu, Zhenyu, et al.
Published: (2025)
by: Liu, Zhenyu, et al.
Published: (2025)
Synonymous Variational Inference for Perceptual Image Compression
by: Liang, Zijian, et al.
Published: (2025)
by: Liang, Zijian, et al.
Published: (2025)
Distance Guided Generative Adversarial Network for Explainable Binary Classifications
by: Xiong, Xiangyu, et al.
Published: (2023)
by: Xiong, Xiangyu, et al.
Published: (2023)
An I2I Inpainting Approach for Efficient Channel Knowledge Map Construction
by: Jin, Zhenzhou, et al.
Published: (2024)
by: Jin, Zhenzhou, et al.
Published: (2024)
Learning a distance measure from the information-estimation geometry of data
by: Ohayon, Guy, et al.
Published: (2025)
by: Ohayon, Guy, et al.
Published: (2025)
A Multi-Drone Multi-View Dataset and Deep Learning Framework for Pedestrian Detection and Tracking
by: Dakic, Kosta, et al.
Published: (2025)
by: Dakic, Kosta, et al.
Published: (2025)
Spherical Leech Quantization for Visual Tokenization and Generation
by: Zhao, Yue, et al.
Published: (2025)
by: Zhao, Yue, et al.
Published: (2025)
Rate-Adaptive Quantization: A Multi-Rate Codebook Adaptation for Vector Quantization-based Generative Models
by: Seo, Jiwan, et al.
Published: (2024)
by: Seo, Jiwan, et al.
Published: (2024)
MedMamba: Vision Mamba for Medical Image Classification
by: Yue, Yubiao, et al.
Published: (2024)
by: Yue, Yubiao, et al.
Published: (2024)
Automated Detection of Myopic Maculopathy in MMAC 2023: Achievements in Classification, Segmentation, and Spherical Equivalent Prediction
by: Li, Yihao, et al.
Published: (2024)
by: Li, Yihao, et al.
Published: (2024)
Toward Lightweight and Fast Decoders for Diffusion Models in Image and Video Generation
by: Buzovkin, Alexey, et al.
Published: (2025)
by: Buzovkin, Alexey, et al.
Published: (2025)
Image Motion Blur Removal in the Temporal Dimension with Video Diffusion Models
by: Pang, Wang, et al.
Published: (2025)
by: Pang, Wang, et al.
Published: (2025)
SSUMamba: Spatial-Spectral Selective State Space Model for Hyperspectral Image Denoising
by: Fu, Guanyiman, et al.
Published: (2024)
by: Fu, Guanyiman, et al.
Published: (2024)
Contrastive Learning and Adversarial Disentanglement for Privacy-Aware Task-Oriented Semantic Communication
by: Erak, Omar, et al.
Published: (2024)
by: Erak, Omar, et al.
Published: (2024)
Explanations of Classifiers Enhance Medical Image Segmentation via End-to-end Pre-training
by: Chen, Jiamin, et al.
Published: (2024)
by: Chen, Jiamin, et al.
Published: (2024)
Aligning Task- and Reconstruction-Oriented Communications for Edge Intelligence
by: Diao, Yufeng, et al.
Published: (2025)
by: Diao, Yufeng, et al.
Published: (2025)
Task-Oriented Co-Design of Communication, Computing, and Control for Edge-Enabled Industrial Cyber-Physical Systems
by: Diao, Yufeng, et al.
Published: (2025)
by: Diao, Yufeng, et al.
Published: (2025)
Investigating Self-Supervised Image Denoising with Denaturation
by: Waida, Hiroki, et al.
Published: (2024)
by: Waida, Hiroki, et al.
Published: (2024)
Generative Video Semantic Communication via Multimodal Semantic Fusion with Large Model
by: Yin, Hang, et al.
Published: (2025)
by: Yin, Hang, et al.
Published: (2025)
GCtx-UNet: Efficient Network for Medical Image Segmentation
by: Alrfou, Khaled, et al.
Published: (2024)
by: Alrfou, Khaled, et al.
Published: (2024)
Revisiting Generative Adversarial Networks for Binary Semantic Segmentation on Imbalanced Datasets
by: Xu, Lei, et al.
Published: (2024)
by: Xu, Lei, et al.
Published: (2024)
Deep Learning Superresolution for 7T Knee MR Imaging: Impact on Image Quality and Diagnostic Performance
by: Chen, Pinzhen, et al.
Published: (2026)
by: Chen, Pinzhen, et al.
Published: (2026)
Translation-based Video-to-Video Synthesis
by: Saha, Pratim, et al.
Published: (2024)
by: Saha, Pratim, et al.
Published: (2024)
Video Quality Enhancement Using Deep Learning-Based Prediction Models for Quantized DCT Coefficients in MPEG I-frames
by: Busson, Antonio J G, et al.
Published: (2020)
by: Busson, Antonio J G, et al.
Published: (2020)
Principled Probabilistic Imaging using Diffusion Models as Plug-and-Play Priors
by: Wu, Zihui, et al.
Published: (2024)
by: Wu, Zihui, et al.
Published: (2024)
RetinaRegen: A Hybrid Model for Readability and Detail Restoration in Fundus Images
by: Tang, Yuhan, et al.
Published: (2025)
by: Tang, Yuhan, et al.
Published: (2025)
UniCompress: Enhancing Multi-Data Medical Image Compression with Knowledge Distillation
by: Yang, Runzhao, et al.
Published: (2024)
by: Yang, Runzhao, et al.
Published: (2024)
Diffusion-Aided Joint Source Channel Coding For High Realism Wireless Image Transmission
by: Yang, Mingyu, et al.
Published: (2024)
by: Yang, Mingyu, et al.
Published: (2024)
ROI-based Deep Image Compression with Implicit Bit Allocation
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
RAGE for the Machine: Image Compression with Low-Cost Random Access for Embedded Applications
by: Rask, Christian D., et al.
Published: (2024)
by: Rask, Christian D., et al.
Published: (2024)
Segment Anything Model for Medical Images?
by: Huang, Yuhao, et al.
Published: (2023)
by: Huang, Yuhao, et al.
Published: (2023)
Implicit Image-to-Image Schrodinger Bridge for Image Restoration
by: Wang, Yuang, et al.
Published: (2024)
by: Wang, Yuang, et al.
Published: (2024)
Video Denoising in Fluorescence Guided Surgery
by: Seets, Trevor, et al.
Published: (2024)
by: Seets, Trevor, et al.
Published: (2024)
MambaVC: Learned Visual Compression with Selective State Spaces
by: Qin, Shiyu, et al.
Published: (2024)
by: Qin, Shiyu, et al.
Published: (2024)
Similar Items
-
GANCompress: GAN-Enhanced Neural Image Compression with Binary Spherical Quantization
by: Sivakoti, Karthik
Published: (2025) -
Bagged Deep Image Prior for Recovering Images in the Presence of Speckle Noise
by: Chen, Xi, et al.
Published: (2024) -
High Perceptual Quality Wireless Image Delivery with Denoising Diffusion Models
by: Yilmaz, Selim F., et al.
Published: (2023) -
Quantifying Knowledge Distillation Using Partial Information Decomposition
by: Dissanayake, Pasan, et al.
Published: (2024) -
Fast and Accurate Cooperative Radio Map Estimation Enabled by GAN
by: Zhang, Zezhong, et al.
Published: (2024)