Latent Denoising Makes Good Tokenizers
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Jiawei, Li, Tianhong, Fan, Lijie, Tian, Yonglong, Wang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2024)
by: Fan, Lijie, et al.
Published: (2024)
Denoising Vision Transformers
by: Yang, Jiawei, et al.
Published: (2024)
by: Yang, Jiawei, et al.
Published: (2024)
Back to Basics: Let Denoising Generative Models Denoise
by: Li, Tianhong, et al.
Published: (2025)
by: Li, Tianhong, et al.
Published: (2025)
Autoregressive Image Generation without Vector Quantization
by: Li, Tianhong, et al.
Published: (2024)
by: Li, Tianhong, et al.
Published: (2024)
Representation Fréchet Loss for Visual Generation
by: Yang, Jiawei, et al.
Published: (2026)
by: Yang, Jiawei, et al.
Published: (2026)
Learning Vision from Models Rivals Learning Vision from Data
by: Tian, Yonglong, et al.
Published: (2023)
by: Tian, Yonglong, et al.
Published: (2023)
Image Understanding Makes for A Good Tokenizer for Image Generation
by: Wang, Luting, et al.
Published: (2024)
by: Wang, Luting, et al.
Published: (2024)
Fractal Generative Models
by: Li, Tianhong, et al.
Published: (2025)
by: Li, Tianhong, et al.
Published: (2025)
Denoising Reuse: Exploiting Inter-frame Motion Consistency for Efficient Video Latent Generation
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2025)
by: Fan, Lijie, et al.
Published: (2025)
Dataset Distillers Are Good Label Denoisers In the Wild
by: Cheng, Lechao, et al.
Published: (2024)
by: Cheng, Lechao, et al.
Published: (2024)
Reparo: Loss-Resilient Generative Codec for Video Conferencing
by: Li, Tianhong, et al.
Published: (2023)
by: Li, Tianhong, et al.
Published: (2023)
MedM-VL: What Makes a Good Medical LVLM?
by: Shi, Yiming, et al.
Published: (2025)
by: Shi, Yiming, et al.
Published: (2025)
Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
End-to-End Training for Unified Tokenization and Latent Denoising
by: Duggal, Shivam, et al.
Published: (2026)
by: Duggal, Shivam, et al.
Published: (2026)
Video Generation Models Are Good Latent Reward Models
by: Mi, Xiaoyue, et al.
Published: (2025)
by: Mi, Xiaoyue, et al.
Published: (2025)
Residual Denoising Diffusion Models
by: Liu, Jiawei, et al.
Published: (2023)
by: Liu, Jiawei, et al.
Published: (2023)
VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
by: Bi, Tianci, et al.
Published: (2025)
by: Bi, Tianci, et al.
Published: (2025)
Making Large Vision Language Models to be Good Few-shot Learners
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
What Makes for Good Image Captions?
by: Chen, Delong, et al.
Published: (2024)
by: Chen, Delong, et al.
Published: (2024)
Autoregressive Denoising Score Matching is a Good Video Anomaly Detector
by: Zhang, Hanwen, et al.
Published: (2025)
by: Zhang, Hanwen, et al.
Published: (2025)
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens
by: Xie, Qingsong, et al.
Published: (2025)
by: Xie, Qingsong, et al.
Published: (2025)
One-step Latent-free Image Generation with Pixel Mean Flows
by: Lu, Yiyang, et al.
Published: (2026)
by: Lu, Yiyang, et al.
Published: (2026)
Enhancing End-to-End Autonomous Driving with Latent World Model
by: Li, Yingyan, et al.
Published: (2024)
by: Li, Yingyan, et al.
Published: (2024)
DeTrack: In-model Latent Denoising Learning for Visual Object Tracking
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
Fréchet Denoised Distance: Enhancing Plausibility Evaluation for Generated Designs with Denoising Autoencoder
by: Fan, Jiajie, et al.
Published: (2024)
by: Fan, Jiajie, et al.
Published: (2024)
What Makes for a Good Stereoscopic Image?
by: Tamir, Netanel Y., et al.
Published: (2024)
by: Tamir, Netanel Y., et al.
Published: (2024)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Latent Denoising Improves Visual Alignment in Large Multimodal Models
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
Denoising Vision Transformer Autoencoder with Spectral Self-Regularization
by: Xiang, Xunzhi, et al.
Published: (2025)
by: Xiang, Xunzhi, et al.
Published: (2025)
Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
by: Liu, Weijia, et al.
Published: (2025)
by: Liu, Weijia, et al.
Published: (2025)
What Makes a Good Dataset for Knowledge Distillation?
by: Frank, Logan, et al.
Published: (2024)
by: Frank, Logan, et al.
Published: (2024)
Return of Unconditional Generation: A Self-supervised Representation Generation Method
by: Li, Tianhong, et al.
Published: (2023)
by: Li, Tianhong, et al.
Published: (2023)
Latent Denoising Diffusion GAN: Faster sampling, Higher image quality
by: Trinh, Luan Thanh, et al.
Published: (2024)
by: Trinh, Luan Thanh, et al.
Published: (2024)
Good Seed Makes a Good Crop: Discovering Secret Seeds in Text-to-Image Diffusion Models
by: Xu, Katherine, et al.
Published: (2024)
by: Xu, Katherine, et al.
Published: (2024)
TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens
by: Ren, Jiawei, et al.
Published: (2026)
by: Ren, Jiawei, et al.
Published: (2026)
Diffusion-Sharpening: Fine-tuning Diffusion Models with Denoising Trajectory Sharpening
by: Tian, Ye, et al.
Published: (2025)
by: Tian, Ye, et al.
Published: (2025)
TAFG-MAN: Timestep-Adaptive Frequency-Gated Latent Diffusion for Efficient and High-Quality Low-Dose CT Image Denoising
by: Fang, Tangtangfang, et al.
Published: (2026)
by: Fang, Tangtangfang, et al.
Published: (2026)
VideoRoPE: What Makes for Good Video Rotary Position Embedding?
by: Wei, Xilin, et al.
Published: (2025)
by: Wei, Xilin, et al.
Published: (2025)
Similar Items
-
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2024) -
Denoising Vision Transformers
by: Yang, Jiawei, et al.
Published: (2024) -
Back to Basics: Let Denoising Generative Models Denoise
by: Li, Tianhong, et al.
Published: (2025) -
Autoregressive Image Generation without Vector Quantization
by: Li, Tianhong, et al.
Published: (2024) -
Representation Fréchet Loss for Visual Generation
by: Yang, Jiawei, et al.
Published: (2026)