Rethinking VLM Representation for VLA Initialization
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Weifeng, Huang, Siyuan, Li, Hao, Chen, Tingwei, An, Ruichuan, Wei, Xinyu, Liu, Jianbo, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
by: Lin, Weifeng, et al.
Published: (2025)
by: Lin, Weifeng, et al.
Published: (2025)
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
by: Li, Muyang, et al.
Published: (2026)
by: Li, Muyang, et al.
Published: (2026)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving
by: Huang, Yuzhou, et al.
Published: (2026)
by: Huang, Yuzhou, et al.
Published: (2026)
ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Driving
by: Peng, Qihang, et al.
Published: (2025)
by: Peng, Qihang, et al.
Published: (2025)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
by: Zhang, Jianke, et al.
Published: (2026)
by: Zhang, Jianke, et al.
Published: (2026)
Exploring Image Representation with Decoupled Classical Visual Descriptors
by: Qu, Chenyuan, et al.
Published: (2025)
by: Qu, Chenyuan, et al.
Published: (2025)
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
by: Xue, Xizhe, et al.
Published: (2024)
by: Xue, Xizhe, et al.
Published: (2024)
An Exploratory Study on Abstract Images and Visual Representations Learned from Them
by: Li, Haotian, et al.
Published: (2025)
by: Li, Haotian, et al.
Published: (2025)
OViP: Online Vision-Language Preference Learning for VLM Hallucination
by: Liu, Shujun, et al.
Published: (2025)
by: Liu, Shujun, et al.
Published: (2025)
Reading Relevant Feature from Global Representation Memory for Visual Object Tracking
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
DirectTriGS: Triplane-based Gaussian Splatting Field Representation for 3D Generation
by: Ju, Xiaoliang, et al.
Published: (2025)
by: Ju, Xiaoliang, et al.
Published: (2025)
Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
by: Zhao, Pu, et al.
Published: (2025)
by: Zhao, Pu, et al.
Published: (2025)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
by: Zhang, Lingfeng, et al.
Published: (2025)
by: Zhang, Lingfeng, et al.
Published: (2025)
VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual Staining
by: Chen, Zizhi, et al.
Published: (2025)
by: Chen, Zizhi, et al.
Published: (2025)
Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution Detection
by: Li, Jinglun, et al.
Published: (2024)
by: Li, Jinglun, et al.
Published: (2024)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
by: Chen, Xinyan, et al.
Published: (2025)
by: Chen, Xinyan, et al.
Published: (2025)
SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
by: Dai, Gaole, et al.
Published: (2025)
by: Dai, Gaole, et al.
Published: (2025)
Rethinking Prior Information Generation with CLIP for Few-Shot Segmentation
by: Wang, Jin, et al.
Published: (2024)
by: Wang, Jin, et al.
Published: (2024)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
by: Han, ByungOk, et al.
Published: (2024)
by: Han, ByungOk, et al.
Published: (2024)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
by: Guo, Wenxuan, et al.
Published: (2026)
by: Guo, Wenxuan, et al.
Published: (2026)
Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and Relabeling
by: Cui, Xiao, et al.
Published: (2025)
by: Cui, Xiao, et al.
Published: (2025)
DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
by: Yin, Bowen, et al.
Published: (2023)
by: Yin, Bowen, et al.
Published: (2023)
IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models
by: Shi, Liang, et al.
Published: (2026)
by: Shi, Liang, et al.
Published: (2026)
Towards Pixel-Level VLM Perception via Simple Points Prediction
by: Song, Tianhui, et al.
Published: (2026)
by: Song, Tianhui, et al.
Published: (2026)
AdaFV: Rethinking of Visual-Language alignment for VLM acceleration
by: Han, Jiayi, et al.
Published: (2025)
by: Han, Jiayi, et al.
Published: (2025)
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
by: Li, Xinze, et al.
Published: (2026)
by: Li, Xinze, et al.
Published: (2026)
VLM-in-the-Loop: A Plug-In Quality Assurance Module for ECG Digitization Pipelines
by: Li, Jiachen, et al.
Published: (2026)
by: Li, Jiachen, et al.
Published: (2026)
Rethinking MLLM Itself as a Segmenter with a Single Segmentation Token
by: Zhang, Anqi, et al.
Published: (2026)
by: Zhang, Anqi, et al.
Published: (2026)
TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center Learning
by: Li, Jinglun, et al.
Published: (2024)
by: Li, Jinglun, et al.
Published: (2024)
Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
by: Wei, Zhihua, et al.
Published: (2026)
by: Wei, Zhihua, et al.
Published: (2026)
AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
by: Zheng, Dian, et al.
Published: (2025)
by: Zheng, Dian, et al.
Published: (2025)
DF-SLAM: Dictionary Factors Representation for High-Fidelity Neural Implicit Dense Visual SLAM System
by: Wei, Weifeng, et al.
Published: (2024)
by: Wei, Weifeng, et al.
Published: (2024)
HDRFace: Rethinking Face Restoration with High-Dimensional Representation
by: Wang, Zirui, et al.
Published: (2026)
by: Wang, Zirui, et al.
Published: (2026)
Segment Anything Is Not Always Perfect: An Investigation of SAM on Different Real-world Applications
by: Ji, Wei, et al.
Published: (2023)
by: Ji, Wei, et al.
Published: (2023)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
by: Chen, Yuxiao, et al.
Published: (2026)
by: Chen, Yuxiao, et al.
Published: (2026)
ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
by: Tian, Changyao, et al.
Published: (2023)
by: Tian, Changyao, et al.
Published: (2023)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
Similar Items
-
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
by: Lin, Weifeng, et al.
Published: (2025) -
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
by: Lin, Weifeng, et al.
Published: (2024) -
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
by: Li, Muyang, et al.
Published: (2026) -
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
by: Lin, Weifeng, et al.
Published: (2024) -
MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving
by: Huang, Yuzhou, et al.
Published: (2026)