VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bi, Tianci, Zhang, Xiaoyi, Lu, Yan, Zheng, Nanning |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaVFM: Adaptive Vision Foundation Models for Edge Intelligence via LLM-Guided Execution
by: Zhao, Yiwei, et al.
Published: (2026)
by: Zhao, Yiwei, et al.
Published: (2026)
LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024)
by: Sadat, Seyedmorteza, et al.
Published: (2024)
Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis
by: Bi, Tianci, et al.
Published: (2024)
by: Bi, Tianci, et al.
Published: (2024)
REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
by: Leng, Xingjian, et al.
Published: (2025)
by: Leng, Xingjian, et al.
Published: (2025)
Multimodal Latent Language Modeling with Next-Token Diffusion
by: Sun, Yutao, et al.
Published: (2024)
by: Sun, Yutao, et al.
Published: (2024)
Quantize-then-Rectify: Efficient VQ-VAE Training
by: Zhang, Borui, et al.
Published: (2025)
by: Zhang, Borui, et al.
Published: (2025)
FSOD-VFM: Few-Shot Object Detection with Vision Foundation Models and Graph Diffusion
by: Feng, Chen-Bin, et al.
Published: (2026)
by: Feng, Chen-Bin, et al.
Published: (2026)
Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation
by: Liu, Shengqi, et al.
Published: (2024)
by: Liu, Shengqi, et al.
Published: (2024)
Token Turing Machines are Efficient Vision Models
by: Jajal, Purvish, et al.
Published: (2024)
by: Jajal, Purvish, et al.
Published: (2024)
Contrastive Conditional-Unconditional Alignment for Long-tailed Diffusion Model
by: Chen, Fang, et al.
Published: (2025)
by: Chen, Fang, et al.
Published: (2025)
Breaking through the learning plateaus of in-context learning in Transformer
by: Fu, Jingwen, et al.
Published: (2023)
by: Fu, Jingwen, et al.
Published: (2023)
Canonical Latent Representations in Conditional Diffusion Models
by: Xu, Yitao, et al.
Published: (2025)
by: Xu, Yitao, et al.
Published: (2025)
Vision Foundation Models in Remote Sensing: A Survey
by: Lu, Siqi, et al.
Published: (2024)
by: Lu, Siqi, et al.
Published: (2024)
PathoTune: Adapting Visual Foundation Model to Pathological Specialists
by: Lu, Jiaxuan, et al.
Published: (2024)
by: Lu, Jiaxuan, et al.
Published: (2024)
REED-VAE: RE-Encode Decode Training for Iterative Image Editing with Diffusion Models
by: Almog, Gal, et al.
Published: (2025)
by: Almog, Gal, et al.
Published: (2025)
AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors
by: Fučka, Matic, et al.
Published: (2026)
by: Fučka, Matic, et al.
Published: (2026)
CellVTA: Enhancing Vision Foundation Models for Accurate Cell Segmentation and Classification
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
Covariance-Aware Goodness for Scalable Forward-Forward Learning
by: Jiang, Xiaoyi, et al.
Published: (2026)
by: Jiang, Xiaoyi, et al.
Published: (2026)
Stable Diffusion Models are Secretly Good at Visual In-Context Learning
by: Oorloff, Trevine, et al.
Published: (2025)
by: Oorloff, Trevine, et al.
Published: (2025)
FairDiffusion: Enhancing Equity in Latent Diffusion Models via Fair Bayesian Perturbation
by: Luo, Yan, et al.
Published: (2024)
by: Luo, Yan, et al.
Published: (2024)
CASL: Concept-Aligned Sparse Latents for Interpreting Diffusion Models
by: He, Zhenghao, et al.
Published: (2026)
by: He, Zhenghao, et al.
Published: (2026)
CAD-VAE: Leveraging Correlation-Aware Latents for Comprehensive Fair Disentanglement
by: Ma, Chenrui, et al.
Published: (2025)
by: Ma, Chenrui, et al.
Published: (2025)
Improved Anomaly Detection through Conditional Latent Space VAE Ensembles
by: Åström, Oskar, et al.
Published: (2024)
by: Åström, Oskar, et al.
Published: (2024)
What Makes "Good" Distractors for Object Hallucination Evaluation in Large Vision-Language Models?
by: Xie, Ming-Kun, et al.
Published: (2025)
by: Xie, Ming-Kun, et al.
Published: (2025)
VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
Exploring Token Pruning in Vision State Space Models
by: Zhan, Zheng, et al.
Published: (2024)
by: Zhan, Zheng, et al.
Published: (2024)
SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation
by: Xing, Ximing, et al.
Published: (2024)
by: Xing, Ximing, et al.
Published: (2024)
VFM-ISRefiner: Towards Better Adapting Vision Foundation Models for Interactive Segmentation of Remote Sensing Images
by: Wang, Deliang, et al.
Published: (2025)
by: Wang, Deliang, et al.
Published: (2025)
Closed-Loop Unsupervised Representation Disentanglement with $β$-VAE Distillation and Diffusion Probabilistic Feedback
by: Jin, Xin, et al.
Published: (2024)
by: Jin, Xin, et al.
Published: (2024)
Can OOD Object Detectors Learn from Foundation Models?
by: Liu, Jiahui, et al.
Published: (2024)
by: Liu, Jiahui, et al.
Published: (2024)
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
KeyPointDiffuser: Unsupervised 3D Keypoint Learning via Latent Diffusion Models
by: Newbury, Rhys, et al.
Published: (2025)
by: Newbury, Rhys, et al.
Published: (2025)
Learning Latent Space Hierarchical EBM Diffusion Models
by: Cui, Jiali, et al.
Published: (2024)
by: Cui, Jiali, et al.
Published: (2024)
Gradient-free Decoder Inversion in Latent Diffusion Models
by: Hong, Seongmin, et al.
Published: (2024)
by: Hong, Seongmin, et al.
Published: (2024)
LURE: Latent Space Unblocking for Multi-Concept Reawakening in Diffusion Models
by: Sun, Mengyu, et al.
Published: (2026)
by: Sun, Mengyu, et al.
Published: (2026)
Hard Cases Detection in Motion Prediction by Vision-Language Foundation Models
by: Yang, Yi, et al.
Published: (2024)
by: Yang, Yi, et al.
Published: (2024)
Revisiting Active Learning in the Era of Vision Foundation Models
by: Gupte, Sanket Rajan, et al.
Published: (2024)
by: Gupte, Sanket Rajan, et al.
Published: (2024)
Exploring Aleatoric Uncertainty in Object Detection via Vision Foundation Models
by: Cui, Peng, et al.
Published: (2024)
by: Cui, Peng, et al.
Published: (2024)
Tutorial on Diffusion Models for Imaging and Vision
by: Chan, Stanley H.
Published: (2024)
by: Chan, Stanley H.
Published: (2024)
TLDR: Token-Level Detective Reward Model for Large Vision Language Models
by: Fu, Deqing, et al.
Published: (2024)
by: Fu, Deqing, et al.
Published: (2024)
Similar Items
-
AdaVFM: Adaptive Vision Foundation Models for Edge Intelligence via LLM-Guided Execution
by: Zhao, Yiwei, et al.
Published: (2026) -
LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024) -
Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis
by: Bi, Tianci, et al.
Published: (2024) -
REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
by: Leng, Xingjian, et al.
Published: (2025) -
Multimodal Latent Language Modeling with Next-Token Diffusion
by: Sun, Yutao, et al.
Published: (2024)