SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Hao, Wang, Ze, Li, Xiang, Sun, Ximeng, Chen, Fangyi, Liu, Jiang, Wang, Jindong, Raj, Bhiksha, Liu, Zicheng, Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
Latent Visual Reasoning
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
Learning from Online Videos at Inference Time for Computer-Use Agents
von: Liu, Yujian, et al.
Veröffentlicht: (2025)
von: Liu, Yujian, et al.
Veröffentlicht: (2025)
Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
von: Wang, Ze, et al.
Veröffentlicht: (2025)
von: Wang, Ze, et al.
Veröffentlicht: (2025)
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
von: Guo, Yuxiang, et al.
Veröffentlicht: (2025)
von: Guo, Yuxiang, et al.
Veröffentlicht: (2025)
DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning
von: Huang, Chao, et al.
Veröffentlicht: (2025)
von: Huang, Chao, et al.
Veröffentlicht: (2025)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
von: Lin, Jingyang, et al.
Veröffentlicht: (2026)
von: Lin, Jingyang, et al.
Veröffentlicht: (2026)
KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
MOVi: Training-free Text-conditioned Multi-Object Video Generation
von: Rahman, Aimon, et al.
Veröffentlicht: (2025)
von: Rahman, Aimon, et al.
Veröffentlicht: (2025)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
XQ-GAN: An Open-source Image Tokenization Framework for Autoregressive Generation
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
CaptionQA: Is Your Caption as Useful as the Image Itself?
von: Yang, Shijia, et al.
Veröffentlicht: (2025)
von: Yang, Shijia, et al.
Veröffentlicht: (2025)
MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
Quantize-then-Rectify: Efficient VQ-VAE Training
von: Zhang, Borui, et al.
Veröffentlicht: (2025)
von: Zhang, Borui, et al.
Veröffentlicht: (2025)
Attentive VQ-VAE
von: Hoyos, Angello, et al.
Veröffentlicht: (2023)
von: Hoyos, Angello, et al.
Veröffentlicht: (2023)
CD4LM: Consistency Distillation and aDaptive Decoding for Diffusion Language Models
von: Liang, Yihao, et al.
Veröffentlicht: (2026)
von: Liang, Yihao, et al.
Veröffentlicht: (2026)
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
von: Singh, Shivam, et al.
Veröffentlicht: (2026)
von: Singh, Shivam, et al.
Veröffentlicht: (2026)
ImageFolder: Autoregressive Image Generation with Folded Tokens
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
DiffSparse: Accelerating Diffusion Transformers with Learned Token Sparsity
von: Zhu, Haowei, et al.
Veröffentlicht: (2026)
von: Zhu, Haowei, et al.
Veröffentlicht: (2026)
Completing Visual Objects via Bridging Generation and Segmentation
von: Li, Xiang, et al.
Veröffentlicht: (2023)
von: Li, Xiang, et al.
Veröffentlicht: (2023)
On Catastrophic Inheritance of Large Foundation Models
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Agent Laboratory: Using LLM Agents as Research Assistants
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2025)
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2025)
Self-Taught Agentic Long Context Understanding
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
Exploring VQ-VAE with Prosody Parameters for Speaker Anonymization
von: Leang, Sotheara, et al.
Veröffentlicht: (2024)
von: Leang, Sotheara, et al.
Veröffentlicht: (2024)
MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
An Embarrassingly Simple Baseline for Imbalanced Semi-Supervised Learning
von: Chen, Hao, et al.
Veröffentlicht: (2022)
von: Chen, Hao, et al.
Veröffentlicht: (2022)
Slight Corruption in Pre-training Data Makes Better Diffusion Models
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Image Tokenizer Needs Post-Training
von: Qiu, Kai, et al.
Veröffentlicht: (2025)
von: Qiu, Kai, et al.
Veröffentlicht: (2025)
Impact of Noisy Supervision in Foundation Model Learning
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
von: Chen, Hao, et al.
Veröffentlicht: (2023)
von: Chen, Hao, et al.
Veröffentlicht: (2023)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
von: Wang, Yating, et al.
Veröffentlicht: (2025)
von: Wang, Yating, et al.
Veröffentlicht: (2025)
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
von: Liu, Huaize, et al.
Veröffentlicht: (2025)
von: Liu, Huaize, et al.
Veröffentlicht: (2025)
Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label Configurations
von: Chen, Hao, et al.
Veröffentlicht: (2023)
von: Chen, Hao, et al.
Veröffentlicht: (2023)
Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEs
von: Cui, Qinpeng, et al.
Veröffentlicht: (2024)
von: Cui, Qinpeng, et al.
Veröffentlicht: (2024)
Class-Partitioned VQ-VAE and Latent Flow Matching for Point Cloud Scene Generation
von: Edirimuni, Dasith de Silva, et al.
Veröffentlicht: (2026)
von: Edirimuni, Dasith de Silva, et al.
Veröffentlicht: (2026)
LADDER: An Efficient Framework for Video Frame Interpolation
von: Shen, Tong, et al.
Veröffentlicht: (2024)
von: Shen, Tong, et al.
Veröffentlicht: (2024)
On Fairness of Unified Multimodal Large Language Model for Image Generation
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
von: li, Bonan, et al.
Veröffentlicht: (2025)
von: li, Bonan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
von: Chen, Hao, et al.
Veröffentlicht: (2025) -
Latent Visual Reasoning
von: Li, Bangzheng, et al.
Veröffentlicht: (2025) -
Learning from Online Videos at Inference Time for Computer-Use Agents
von: Liu, Yujian, et al.
Veröffentlicht: (2025) -
Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
von: Wang, Ze, et al.
Veröffentlicht: (2025) -
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
von: Guo, Yuxiang, et al.
Veröffentlicht: (2025)