Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Shenghao, Zu, Wenqiang, Zhao, Mingyang, Su, Duo, Liu, Shilong, Shi, Ruohua, Li, Guoqi, Zhang, Shanghang, Ma, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Embedded Visual Prompt Tuning
by: Zu, Wenqiang, et al.
Published: (2024)
by: Zu, Wenqiang, et al.
Published: (2024)
Training-Free Representation Guidance for Diffusion Models with a Representation Alignment Projector
by: Zu, Wenqiang, et al.
Published: (2026)
by: Zu, Wenqiang, et al.
Published: (2026)
Pre-trained Models Succeed in Medical Imaging with Representation Similarity Degradation
by: Zu, Wenqiang, et al.
Published: (2025)
by: Zu, Wenqiang, et al.
Published: (2025)
NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
by: Ye, Junliang, et al.
Published: (2025)
by: Ye, Junliang, et al.
Published: (2025)
Exploring Representation Invariance in Finetuning
by: Zu, Wenqiang, et al.
Published: (2025)
by: Zu, Wenqiang, et al.
Published: (2025)
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
by: Jiang, Zhiyuan, et al.
Published: (2025)
by: Jiang, Zhiyuan, et al.
Published: (2025)
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
by: Feng, Duanyu, et al.
Published: (2024)
by: Feng, Duanyu, et al.
Published: (2024)
AI Security in the Foundation Model Era: A Comprehensive Survey from a Unified Perspective
by: Wang, Zhenyi, et al.
Published: (2026)
by: Wang, Zhenyi, et al.
Published: (2026)
From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Models
by: Huang, Kung-Hsiang, et al.
Published: (2024)
by: Huang, Kung-Hsiang, et al.
Published: (2024)
ShapeMamba-EM: Fine-Tuning Foundation Model with Local Shape Descriptors and Mamba Blocks for 3D EM Image Segmentation
by: Shi, Ruohua, et al.
Published: (2024)
by: Shi, Ruohua, et al.
Published: (2024)
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
by: Bai, Shuanghao, et al.
Published: (2025)
by: Bai, Shuanghao, et al.
Published: (2025)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
by: Ma, Yiyang, et al.
Published: (2024)
by: Ma, Yiyang, et al.
Published: (2024)
Autoregressive Image Generation with Randomized Parallel Decoding
by: Li, Haopeng, et al.
Published: (2025)
by: Li, Haopeng, et al.
Published: (2025)
Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
by: Ma, Xiaoxiao, et al.
Published: (2025)
by: Ma, Xiaoxiao, et al.
Published: (2025)
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
by: Fu, Shenghao, et al.
Published: (2024)
by: Fu, Shenghao, et al.
Published: (2024)
UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation
by: Zhang, Ruiheng, et al.
Published: (2026)
by: Zhang, Ruiheng, et al.
Published: (2026)
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
by: Bai, Shuanghao, et al.
Published: (2025)
by: Bai, Shuanghao, et al.
Published: (2025)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025)
by: Zheng, Anlin, et al.
Published: (2025)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2025)
by: Fan, Lijie, et al.
Published: (2025)
Towards Sampling Data Structures for Tensor Products in Turnstile Streams
by: Song, Zhao, et al.
Published: (2025)
by: Song, Zhao, et al.
Published: (2025)
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
by: Liu, Jiaming, et al.
Published: (2025)
by: Liu, Jiaming, et al.
Published: (2025)
Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
by: Wang, Peiyu, et al.
Published: (2025)
by: Wang, Peiyu, et al.
Published: (2025)
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
by: Zhuang, Xianwei, et al.
Published: (2025)
by: Zhuang, Xianwei, et al.
Published: (2025)
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
by: Ye, Junliang, et al.
Published: (2025)
by: Ye, Junliang, et al.
Published: (2025)
Scalable Autoregressive Image Generation with Mamba
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding
by: Ma, Chenglong, et al.
Published: (2025)
by: Ma, Chenglong, et al.
Published: (2025)
A Survey on Image Quality Assessment: Insights, Analysis, and Future Outlook
by: Ma, Chengqian, et al.
Published: (2025)
by: Ma, Chengqian, et al.
Published: (2025)
Autoregressive Meta-Actions for Unified Controllable Trajectory Generation
by: Zhao, Jianbo, et al.
Published: (2025)
by: Zhao, Jianbo, et al.
Published: (2025)
Performance Bounds and Degree-Distribution Optimization of Finite-Length BATS Codes
by: Zhu, Mingyang, et al.
Published: (2025)
by: Zhu, Mingyang, et al.
Published: (2025)
ObjEmbed: Towards Universal Multimodal Object Embeddings
by: Fu, Shenghao, et al.
Published: (2026)
by: Fu, Shenghao, et al.
Published: (2026)
Autoregressive Models in Vision: A Survey
by: Xiong, Jing, et al.
Published: (2024)
by: Xiong, Jing, et al.
Published: (2024)
A Survey on Vision Autoregressive Model
by: Jiang, Kai, et al.
Published: (2024)
by: Jiang, Kai, et al.
Published: (2024)
A New Era in Computational Pathology: A Survey on Foundation and Vision-Language Models
by: Chanda, Dibaloke, et al.
Published: (2024)
by: Chanda, Dibaloke, et al.
Published: (2024)
Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
In-Place Panoptic Radiance Field Segmentation with Perceptual Prior for 3D Scene Understanding
by: Li, Shenghao
Published: (2024)
by: Li, Shenghao
Published: (2024)
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
by: Shen, Yang, et al.
Published: (2024)
by: Shen, Yang, et al.
Published: (2024)
Unified Cross-Scale 3D Generation and Understanding via Autoregressive Modeling
by: Lu, Shuqi, et al.
Published: (2025)
by: Lu, Shuqi, et al.
Published: (2025)
One for All: Toward Unified Foundation Models for Earth Vision
by: Xiong, Zhitong, et al.
Published: (2024)
by: Xiong, Zhitong, et al.
Published: (2024)
Towards a Unified Copernicus Foundation Model for Earth Vision
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Similar Items
-
Embedded Visual Prompt Tuning
by: Zu, Wenqiang, et al.
Published: (2024) -
Training-Free Representation Guidance for Diffusion Models with a Representation Alignment Projector
by: Zu, Wenqiang, et al.
Published: (2026) -
Pre-trained Models Succeed in Medical Imaging with Representation Similarity Degradation
by: Zu, Wenqiang, et al.
Published: (2025) -
NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
by: Ye, Junliang, et al.
Published: (2025) -
Exploring Representation Invariance in Finetuning
by: Zu, Wenqiang, et al.
Published: (2025)