Vision Foundation Models as Generalist Tokenizers for Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Anlin, Han, Qi, Wen, Xin, Ma, Chuofan, Gong, Lanxi, Yu, Gang, Zhang, Xiangyu, Qi, Xiaojuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025)
by: Zheng, Anlin, et al.
Published: (2025)
Hita: Holistic Tokenizer for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025)
by: Zheng, Anlin, et al.
Published: (2025)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
by: Ma, Chuofan, et al.
Published: (2025)
by: Ma, Chuofan, et al.
Published: (2025)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
by: Ma, Chuofan, et al.
Published: (2024)
by: Ma, Chuofan, et al.
Published: (2024)
Learning from Neighbors: Category Extrapolation for Long-Tail Learning
by: Zhao, Shizhen, et al.
Published: (2024)
by: Zhao, Shizhen, et al.
Published: (2024)
Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
by: Zhao, Shizhen, et al.
Published: (2025)
by: Zhao, Shizhen, et al.
Published: (2025)
Can OOD Object Detectors Learn from Foundation Models?
by: Liu, Jiahui, et al.
Published: (2024)
by: Liu, Jiahui, et al.
Published: (2024)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens
by: Wang, Yuqing, et al.
Published: (2026)
by: Wang, Yuqing, et al.
Published: (2026)
"Principal Components" Enable A New Language of Images
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
DINO-Tok: Adapting DINO for Visual Tokenizers
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
by: Huang, Xinrui, et al.
Published: (2025)
by: Huang, Xinrui, et al.
Published: (2025)
Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse
by: Zhang, Kuan, et al.
Published: (2026)
by: Zhang, Kuan, et al.
Published: (2026)
Image Generators are Generalist Vision Learners
by: Gabeur, Valentin, et al.
Published: (2026)
by: Gabeur, Valentin, et al.
Published: (2026)
MedVersa: A Generalist Foundation Model for Medical Image Interpretation
by: Zhou, Hong-Yu, et al.
Published: (2024)
by: Zhou, Hong-Yu, et al.
Published: (2024)
Classes Are Not Equal: An Empirical Study on Image Recognition Fairness
by: Cui, Jiequan, et al.
Published: (2024)
by: Cui, Jiequan, et al.
Published: (2024)
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
EyeFound: A Multimodal Generalist Foundation Model for Ophthalmic Imaging
by: Shi, Danli, et al.
Published: (2024)
by: Shi, Danli, et al.
Published: (2024)
Learning A Low-Level Vision Generalist via Visual Task Prompt
by: Chen, Xiangyu, et al.
Published: (2024)
by: Chen, Xiangyu, et al.
Published: (2024)
MapSR: Prompt-Driven Land Cover Map Super-Resolution via Vision Foundation Models
by: Wang, Ruiqi, et al.
Published: (2026)
by: Wang, Ruiqi, et al.
Published: (2026)
TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection
by: Abdullah, Ahmed, et al.
Published: (2026)
by: Abdullah, Ahmed, et al.
Published: (2026)
GO-NeRF: Generating Objects in Neural Radiance Fields for Virtual Reality Content Creation
by: Dai, Peng, et al.
Published: (2024)
by: Dai, Peng, et al.
Published: (2024)
Debiasing Text-to-Image Diffusion Models
by: He, Ruifei, et al.
Published: (2024)
by: He, Ruifei, et al.
Published: (2024)
Balancing Multi-Target Semi-Supervised Medical Image Segmentation with Collaborative Generalist and Specialists
by: Wang, You, et al.
Published: (2025)
by: Wang, You, et al.
Published: (2025)
Training-free Token Reduction for Vision Mamba
by: Ma, Qiankun, et al.
Published: (2025)
by: Ma, Qiankun, et al.
Published: (2025)
BRIGHT: A Collaborative Generalist-Specialist Foundation Model for Breast Pathology
by: Guo, Xiaojing, et al.
Published: (2026)
by: Guo, Xiaojing, et al.
Published: (2026)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
VisionFM: a Multi-Modal Multi-Task Vision Foundation Model for Generalist Ophthalmic Artificial Intelligence
by: Qiu, Jianing, et al.
Published: (2023)
by: Qiu, Jianing, et al.
Published: (2023)
Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic Segmentation
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
by: Bi, Tianci, et al.
Published: (2025)
by: Bi, Tianci, et al.
Published: (2025)
QDM: Quadtree-Based Region-Adaptive Sparse Diffusion Models for Efficient Image Super-Resolution
by: Yang, Donglin, et al.
Published: (2025)
by: Yang, Donglin, et al.
Published: (2025)
Improved Masked Image Generation with Knowledge-Augmented Token Representations
by: Liang, Guotao, et al.
Published: (2025)
by: Liang, Guotao, et al.
Published: (2025)
InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists
by: Gan, Yulu, et al.
Published: (2023)
by: Gan, Yulu, et al.
Published: (2023)
ObjectMorpher: 3D-Aware Image Editing via Deformable 3DGS Models
by: Xie, Yuhuan, et al.
Published: (2026)
by: Xie, Yuhuan, et al.
Published: (2026)
Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models
by: Chen, Zizhi, et al.
Published: (2025)
by: Chen, Zizhi, et al.
Published: (2025)
From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models
by: Gong, Weile, et al.
Published: (2026)
by: Gong, Weile, et al.
Published: (2026)
Vision Generalist Model: A Survey
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Masked AutoDecoder is Effective Multi-Task Vision Generalist
by: Qiu, Han, et al.
Published: (2024)
by: Qiu, Han, et al.
Published: (2024)
EscherNet: A Generative Model for Scalable View Synthesis
by: Kong, Xin, et al.
Published: (2024)
by: Kong, Xin, et al.
Published: (2024)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Similar Items
-
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025) -
Hita: Holistic Tokenizer for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025) -
UniTok: A Unified Tokenizer for Visual Generation and Understanding
by: Ma, Chuofan, et al.
Published: (2025) -
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
by: Ma, Chuofan, et al.
Published: (2024) -
Learning from Neighbors: Category Extrapolation for Long-Tail Learning
by: Zhao, Shizhen, et al.
Published: (2024)