Spherical Leech Quantization for Visual Tokenization and Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Yue, Jiang, Hanwen, Xu, Zhenlin, Yang, Chutong, Adeli, Ehsan, Krähenbühl, Philipp |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Image and Video Tokenization with Binary Spherical Quantization
por: Zhao, Yue, et al.
Publicado: (2024)
por: Zhao, Yue, et al.
Publicado: (2024)
Spectral Vision Transformer for Efficient Tokenization with Limited Data
por: Roberts, Alexandra G., et al.
Publicado: (2026)
por: Roberts, Alexandra G., et al.
Publicado: (2026)
Vector Quantization for Deep-Learning-Based CSI Feedback in Massive MIMO Systems
por: Shin, Junyong, et al.
Publicado: (2024)
por: Shin, Junyong, et al.
Publicado: (2024)
CognitionCapturer: Decoding Visual Stimuli From Human EEG Signal With Multimodal Information
por: Zhang, Kaifan, et al.
Publicado: (2024)
por: Zhang, Kaifan, et al.
Publicado: (2024)
VCHAR:Variance-Driven Complex Human Activity Recognition framework with Generative Representation
por: Sun, Yuan, et al.
Publicado: (2024)
por: Sun, Yuan, et al.
Publicado: (2024)
Modality-Aware and Anatomical Vector-Quantized Autoencoding for Multimodal Brain MRI
por: Li, Mingjie, et al.
Publicado: (2026)
por: Li, Mingjie, et al.
Publicado: (2026)
SynthBrainGrow: Synthetic Diffusion Brain Aging for Longitudinal MRI Data Generation in Young People
por: Zapaishchykova, Anna, et al.
Publicado: (2024)
por: Zapaishchykova, Anna, et al.
Publicado: (2024)
DREAM: Dynamic Retinal Enhancement with Adaptive Multi-modal Fusion for Expert Precision Medical Report Generation
por: Shaik, Nagur Shareef, et al.
Publicado: (2026)
por: Shaik, Nagur Shareef, et al.
Publicado: (2026)
Spatial-Related Sensors Matters: 3D Human Motion Reconstruction Assisted with Textual Semantics
por: Yang, Xueyuan, et al.
Publicado: (2023)
por: Yang, Xueyuan, et al.
Publicado: (2023)
Large Language Model-Driven Distributed Integrated Multimodal Sensing and Semantic Communications
por: Peng, Yubo, et al.
Publicado: (2025)
por: Peng, Yubo, et al.
Publicado: (2025)
KNN-MMD: Cross Domain Wireless Sensing via Local Distribution Alignment
por: Zhao, Zijian, et al.
Publicado: (2024)
por: Zhao, Zijian, et al.
Publicado: (2024)
RT-Pose: A 4D Radar Tensor-based 3D Human Pose Estimation and Localization Benchmark
por: Ho, Yuan-Hao, et al.
Publicado: (2024)
por: Ho, Yuan-Hao, et al.
Publicado: (2024)
GenHPE: Generative Counterfactuals for 3D Human Pose Estimation with Radio Frequency Signals
por: Huang, Shuokang, et al.
Publicado: (2025)
por: Huang, Shuokang, et al.
Publicado: (2025)
Accurate Patient Alignment without Unnecessary Imaging Dose via Synthesizing Patient-specific 3D CT Images from 2D kV Images
por: Ding, Yuzhen, et al.
Publicado: (2024)
por: Ding, Yuzhen, et al.
Publicado: (2024)
Multi-View Hierarchical Representation Learning of Fetal Hemodynamics for Maternal Hypertension Detection at the Edge
por: Rafiei, Alireza, et al.
Publicado: (2026)
por: Rafiei, Alireza, et al.
Publicado: (2026)
CaFNet: A Confidence-Driven Framework for Radar Camera Depth Estimation
por: Sun, Huawei, et al.
Publicado: (2024)
por: Sun, Huawei, et al.
Publicado: (2024)
Combining SAR Simulators to Train ATR Models with Synthetic Data
por: Camus, Benjamin, et al.
Publicado: (2025)
por: Camus, Benjamin, et al.
Publicado: (2025)
Human Presence Detection via Wi-Fi Range-Filtered Doppler Spectrum on Commodity Laptops
por: Sanson, Jessica, et al.
Publicado: (2026)
por: Sanson, Jessica, et al.
Publicado: (2026)
Towards Cognitive Defect Analysis in Active Infrared Thermography with Vision-Text Cues
por: Salah, Mohammed, et al.
Publicado: (2026)
por: Salah, Mohammed, et al.
Publicado: (2026)
Evaluating Feature Attribution Methods for Electrocardiogram
por: Suh, Jangwon, et al.
Publicado: (2022)
por: Suh, Jangwon, et al.
Publicado: (2022)
Efficient 4D Radar Data Auto-labeling Method using LiDAR-based Object Detection Network
por: Sun, Min-Hyeok, et al.
Publicado: (2024)
por: Sun, Min-Hyeok, et al.
Publicado: (2024)
Region-Affinity Attention for Whole-Slide Breast Cancer Classification in Deep Ultraviolet Imaging
por: Shaik, Nagur Shareef, et al.
Publicado: (2026)
por: Shaik, Nagur Shareef, et al.
Publicado: (2026)
Process Optimization and Deployment for Sensor-Based Human Activity Recognition Based on Deep Learning
por: Liu, Hanyu, et al.
Publicado: (2025)
por: Liu, Hanyu, et al.
Publicado: (2025)
Hybrid operator learning of wave scattering maps in high-contrast media
por: Balaji, Advait, et al.
Publicado: (2026)
por: Balaji, Advait, et al.
Publicado: (2026)
Adaptive Signal Analysis for Automated Subsurface Defect Detection Using Impact Echo in Concrete Slabs
por: Pavurala, Deepthi, et al.
Publicado: (2024)
por: Pavurala, Deepthi, et al.
Publicado: (2024)
EMind: A Foundation Model for Multi-task Electromagnetic Signals Understanding
por: Luo, Luqing, et al.
Publicado: (2025)
por: Luo, Luqing, et al.
Publicado: (2025)
Real-Time Human Activity Recognition on Edge Microcontrollers: Dynamic Hierarchical Inference with Multi-Spectral Sensor Fusion
por: Li, Boyu, et al.
Publicado: (2026)
por: Li, Boyu, et al.
Publicado: (2026)
RAPTR: Radar-based 3D Pose Estimation using Transformer
por: Kato, Sorachi, et al.
Publicado: (2025)
por: Kato, Sorachi, et al.
Publicado: (2025)
Fusion of Deep Learning and GIS for Advanced Remote Sensing Image Analysis
por: Afroosheh, Sajjad, et al.
Publicado: (2024)
por: Afroosheh, Sajjad, et al.
Publicado: (2024)
GET-UP: GEomeTric-aware Depth Estimation with Radar Points UPsampling
por: Sun, Huawei, et al.
Publicado: (2024)
por: Sun, Huawei, et al.
Publicado: (2024)
Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion
por: Lim, DongHoon, et al.
Publicado: (2025)
por: Lim, DongHoon, et al.
Publicado: (2025)
Scaling Image Tokenizers with Grouped Spherical Quantization
por: Wang, Jiangtao, et al.
Publicado: (2024)
por: Wang, Jiangtao, et al.
Publicado: (2024)
Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment
por: Dai, Shenghong, et al.
Publicado: (2024)
por: Dai, Shenghong, et al.
Publicado: (2024)
Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario Generation
por: Yang, Xiuyu, et al.
Publicado: (2025)
por: Yang, Xiuyu, et al.
Publicado: (2025)
Biometric Authentication Based on Enhanced Remote Photoplethysmography Signal Morphology
por: Sun, Zhaodong, et al.
Publicado: (2024)
por: Sun, Zhaodong, et al.
Publicado: (2024)
A Physics-guided Generative AI Toolkit for Geophysical Monitoring
por: Yang, Junhuan, et al.
Publicado: (2024)
por: Yang, Junhuan, et al.
Publicado: (2024)
Unleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning of Language Models
por: Qin, Yulei, et al.
Publicado: (2024)
por: Qin, Yulei, et al.
Publicado: (2024)
Robust Simultaneous Multislice MRI Reconstruction Using Slice-Wise Learned Generative Diffusion Priors
por: Huang, Shoujin, et al.
Publicado: (2024)
por: Huang, Shoujin, et al.
Publicado: (2024)
Deep Imbalanced Regression to Estimate Vascular Age from PPG Data: a Novel Digital Biomarker for Cardiovascular Health
por: Nie, Guangkun, et al.
Publicado: (2024)
por: Nie, Guangkun, et al.
Publicado: (2024)
Fréchet Power-Scenario Distance: A Metric for Evaluating Generative AI Models across Multiple Time-Scales in Smart Grids
por: Cai, Yuting, et al.
Publicado: (2025)
por: Cai, Yuting, et al.
Publicado: (2025)
Ejemplares similares
-
Image and Video Tokenization with Binary Spherical Quantization
por: Zhao, Yue, et al.
Publicado: (2024) -
Spectral Vision Transformer for Efficient Tokenization with Limited Data
por: Roberts, Alexandra G., et al.
Publicado: (2026) -
Vector Quantization for Deep-Learning-Based CSI Feedback in Massive MIMO Systems
por: Shin, Junyong, et al.
Publicado: (2024) -
CognitionCapturer: Decoding Visual Stimuli From Human EEG Signal With Multimodal Information
por: Zhang, Kaifan, et al.
Publicado: (2024) -
VCHAR:Variance-Driven Complex Human Activity Recognition framework with Generative Representation
por: Sun, Yuan, et al.
Publicado: (2024)