Vision Generalist Model: A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Ziyi, Rao, Yongming, Sun, Shuofeng, Liu, Xinrun, Wei, Yi, Yu, Xumin, Liu, Zuyan, Wang, Yanbo, Liu, Hongmin, Zhou, Jie, Lu, Jiwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
X-3D: Explicit 3D Structure Modeling for Point Cloud Recognition
von: Sun, Shuofeng, et al.
Veröffentlicht: (2024)
von: Sun, Shuofeng, et al.
Veröffentlicht: (2024)
XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation
von: Wang, Ziyi, et al.
Veröffentlicht: (2024)
von: Wang, Ziyi, et al.
Veröffentlicht: (2024)
Efficient Inference of Vision Instruction-Following Models with Elastic Cache
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
Ola: Pushing the Frontiers of Omni-Modal Language Model
von: Liu, Zuyan, et al.
Veröffentlicht: (2025)
von: Liu, Zuyan, et al.
Veröffentlicht: (2025)
OGGSplat: Open Gaussian Growing for Generalizable Reconstruction with Expanded Field-of-View
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
3D Small Object Detection with Dynamic Spatial Pruning
von: Xu, Xiuwei, et al.
Veröffentlicht: (2023)
von: Xu, Xiuwei, et al.
Veröffentlicht: (2023)
Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models
von: Dong, Yuhao, et al.
Veröffentlicht: (2026)
von: Dong, Yuhao, et al.
Veröffentlicht: (2026)
FlowTurbo: Towards Real-time Flow-Based Image Generation with Velocity Refiner
von: Zhao, Wenliang, et al.
Veröffentlicht: (2024)
von: Zhao, Wenliang, et al.
Veröffentlicht: (2024)
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
von: Dong, Yuhao, et al.
Veröffentlicht: (2024)
von: Dong, Yuhao, et al.
Veröffentlicht: (2024)
Length Matters: Length-Aware Transformer for Temporal Sentence Grounding
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
GlobalMamba: Global Image Serialization for Vision Mamba
von: Wang, Chengkun, et al.
Veröffentlicht: (2024)
von: Wang, Chengkun, et al.
Veröffentlicht: (2024)
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
EfficientLLaVA:Generalizable Auto-Pruning for Large Vision-language Models
von: Liang, Yinan, et al.
Veröffentlicht: (2025)
von: Liang, Yinan, et al.
Veröffentlicht: (2025)
Q-VLM: Post-training Quantization for Large Vision-Language Models
von: Wang, Changyuan, et al.
Veröffentlicht: (2024)
von: Wang, Changyuan, et al.
Veröffentlicht: (2024)
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents
von: X, Tencent Robotics, et al.
Veröffentlicht: (2026)
von: X, Tencent Robotics, et al.
Veröffentlicht: (2026)
What Matters in Building Vision-Language-Action Models for Generalist Robots
von: Li, Xinghang, et al.
Veröffentlicht: (2024)
von: Li, Xinghang, et al.
Veröffentlicht: (2024)
Joint 3D Geometry Reconstruction and Motion Generation for 4D Synthesis from a Single Image
von: Zhang, Yanran, et al.
Veröffentlicht: (2025)
von: Zhang, Yanran, et al.
Veröffentlicht: (2025)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
von: Wei, Cong, et al.
Veröffentlicht: (2024)
von: Wei, Cong, et al.
Veröffentlicht: (2024)
From Generalist to Specialist: Adapting Vision Language Models via Task-Specific Visual Instruction Tuning
von: Bai, Yang, et al.
Veröffentlicht: (2024)
von: Bai, Yang, et al.
Veröffentlicht: (2024)
Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
Moaw: Unleashing Motion Awareness for Video Diffusion Models
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
Learning A Low-Level Vision Generalist via Visual Task Prompt
von: Chen, Xiangyu, et al.
Veröffentlicht: (2024)
von: Chen, Xiangyu, et al.
Veröffentlicht: (2024)
DC-Solver: Improving Predictor-Corrector Diffusion Sampler via Dynamic Compensation
von: Zhao, Wenliang, et al.
Veröffentlicht: (2024)
von: Zhao, Wenliang, et al.
Veröffentlicht: (2024)
Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking
von: Zhu, Deyi, et al.
Veröffentlicht: (2026)
von: Zhu, Deyi, et al.
Veröffentlicht: (2026)
Faceptor: A Generalist Model for Face Perception
von: Qin, Lixiong, et al.
Veröffentlicht: (2024)
von: Qin, Lixiong, et al.
Veröffentlicht: (2024)
UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
von: Chen, Dengbo, et al.
Veröffentlicht: (2025)
von: Chen, Dengbo, et al.
Veröffentlicht: (2025)
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
von: Yin, Hang, et al.
Veröffentlicht: (2025)
von: Yin, Hang, et al.
Veröffentlicht: (2025)
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
von: Liu, Fangfu, et al.
Veröffentlicht: (2026)
von: Liu, Fangfu, et al.
Veröffentlicht: (2026)
Towards Accurate Post-training Quantization for Diffusion Models
von: Wang, Changyuan, et al.
Veröffentlicht: (2023)
von: Wang, Changyuan, et al.
Veröffentlicht: (2023)
LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
SFTok: Bridging the Performance Gap in Discrete Tokenizers
von: Rao, Qihang, et al.
Veröffentlicht: (2025)
von: Rao, Qihang, et al.
Veröffentlicht: (2025)
Quantize-then-Rectify: Efficient VQ-VAE Training
von: Zhang, Borui, et al.
Veröffentlicht: (2025)
von: Zhang, Borui, et al.
Veröffentlicht: (2025)
Diffusion Model as a Generalist Segmentation Learner
von: Wang, Haoxiao, et al.
Veröffentlicht: (2026)
von: Wang, Haoxiao, et al.
Veröffentlicht: (2026)
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
von: Jiang, Songtao, et al.
Veröffentlicht: (2025)
von: Jiang, Songtao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
von: Liu, Zuyan, et al.
Veröffentlicht: (2024) -
SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs
von: Wang, Jiahui, et al.
Veröffentlicht: (2025) -
X-3D: Explicit 3D Structure Modeling for Point Cloud Recognition
von: Sun, Shuofeng, et al.
Veröffentlicht: (2024) -
XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation
von: Wang, Ziyi, et al.
Veröffentlicht: (2024) -
Efficient Inference of Vision Instruction-Following Models with Elastic Cache
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)