CODA: Repurposing Continuous VAEs for Discrete Tokenization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zeyu, Ni, Zanlin, Hua, Yeguo, Deng, Xin, Ma, Xiao, Zhong, Cheng, Huang, Gao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
by: Liu, Zeyu, et al.
Published: (2026)
by: Liu, Zeyu, et al.
Published: (2026)
ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis
by: Ni, Zanlin, et al.
Published: (2024)
by: Ni, Zanlin, et al.
Published: (2024)
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
by: Zhou, Renping, et al.
Published: (2025)
by: Zhou, Renping, et al.
Published: (2025)
Deep Incomplete Multi-view Learning via Cyclic Permutation of VAEs
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
by: Ni, Zanlin, et al.
Published: (2024)
by: Ni, Zanlin, et al.
Published: (2024)
AdaGen: Learning Adaptive Policy for Image Synthesis
by: Ni, Zanlin, et al.
Published: (2026)
by: Ni, Zanlin, et al.
Published: (2026)
Cross-Modal Adapter for Vision-Language Retrieval
by: Jiang, Haojun, et al.
Published: (2022)
by: Jiang, Haojun, et al.
Published: (2022)
MacTok: Robust Continuous Tokenization for Image Generation
by: Zeng, Hengyu, et al.
Published: (2026)
by: Zeng, Hengyu, et al.
Published: (2026)
LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
by: Xu, Ruyi, et al.
Published: (2024)
by: Xu, Ruyi, et al.
Published: (2024)
GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEs
by: Basioti, Kalliopi, et al.
Published: (2025)
by: Basioti, Kalliopi, et al.
Published: (2025)
FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous Tokens
by: Zhong, Yiming, et al.
Published: (2025)
by: Zhong, Yiming, et al.
Published: (2025)
D-CODA: Diffusion for Coordinated Dual-Arm Data Augmentation
by: Liu, I-Chun Arthur, et al.
Published: (2025)
by: Liu, I-Chun Arthur, et al.
Published: (2025)
On the Adversarial Robustness of Discrete Image Tokenizers
by: Bhagwatkar, Rishika, et al.
Published: (2026)
by: Bhagwatkar, Rishika, et al.
Published: (2026)
Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice
by: Huang, Xiusheng, et al.
Published: (2026)
by: Huang, Xiusheng, et al.
Published: (2026)
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
by: Yue, Yang, et al.
Published: (2026)
by: Yue, Yang, et al.
Published: (2026)
Frequency Autoregressive Image Generation with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
by: Ye, Haotian, et al.
Published: (2025)
by: Ye, Haotian, et al.
Published: (2025)
AdverX-Ray: Ensuring X-Ray Integrity Through Frequency-Sensitive Adversarial VAEs
by: Caetano, Francisco, et al.
Published: (2025)
by: Caetano, Francisco, et al.
Published: (2025)
Adversarial robustness of VAEs through the lens of local geometry
by: Khan, Asif, et al.
Published: (2022)
by: Khan, Asif, et al.
Published: (2022)
Learning Energy-based Variational Latent Prior for VAEs
by: Dutta, Debottam, et al.
Published: (2025)
by: Dutta, Debottam, et al.
Published: (2025)
CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning
by: Sun, Zeyi, et al.
Published: (2025)
by: Sun, Zeyi, et al.
Published: (2025)
DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding
by: Cho, Jungbin, et al.
Published: (2024)
by: Cho, Jungbin, et al.
Published: (2024)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
by: Kim, Dongwon, et al.
Published: (2026)
by: Kim, Dongwon, et al.
Published: (2026)
TMCIR: Token Merge Benefits Composed Image Retrieval
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization
by: Zhang, Yiwei, et al.
Published: (2024)
by: Zhang, Yiwei, et al.
Published: (2024)
Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
by: Wang, Weixing, et al.
Published: (2025)
by: Wang, Weixing, et al.
Published: (2025)
IDPruner: Harmonizing Importance and Diversity in Visual Token Pruning for MLLMs
by: Tan, Yifan, et al.
Published: (2026)
by: Tan, Yifan, et al.
Published: (2026)
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
by: Tan, Zhentao, et al.
Published: (2024)
by: Tan, Zhentao, et al.
Published: (2024)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
by: Zhong, Tianxiong, et al.
Published: (2025)
by: Zhong, Tianxiong, et al.
Published: (2025)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
by: Qu, Liao, et al.
Published: (2024)
by: Qu, Liao, et al.
Published: (2024)
Repurposing Stable Diffusion Attention for Training-Free Unsupervised Interactive Segmentation
by: Karmann, Markus, et al.
Published: (2024)
by: Karmann, Markus, et al.
Published: (2024)
SuperGaussian: Repurposing Video Models for 3D Super Resolution
by: Shen, Yuan, et al.
Published: (2024)
by: Shen, Yuan, et al.
Published: (2024)
ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving
by: Chen, Kai, et al.
Published: (2025)
by: Chen, Kai, et al.
Published: (2025)
Perceptual Group Tokenizer: Building Perception with Iterative Grouping
by: Deng, Zhiwei, et al.
Published: (2023)
by: Deng, Zhiwei, et al.
Published: (2023)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
by: Ma, Chuofan, et al.
Published: (2025)
by: Ma, Chuofan, et al.
Published: (2025)
Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs
by: Lin, Yuhui, et al.
Published: (2026)
by: Lin, Yuhui, et al.
Published: (2026)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025)
by: Zheng, Anlin, et al.
Published: (2025)
Similar Items
-
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
by: Liu, Zeyu, et al.
Published: (2026) -
ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis
by: Ni, Zanlin, et al.
Published: (2024) -
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
by: Zhou, Renping, et al.
Published: (2025) -
Deep Incomplete Multi-view Learning via Cyclic Permutation of VAEs
by: Gao, Xin, et al.
Published: (2025) -
Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
by: Ni, Zanlin, et al.
Published: (2024)