Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Ganlin, Zhang, Tianyi, Hao, Haoran, Wang, Weiyun, Liu, Yibin, Wang, Dehui, Chen, Guanzhou, Cai, Zijian, Chen, Junting, Su, Weijie, Zhou, Wengang, Qiao, Yu, Dai, Jifeng, Pang, Jiangmiao, Luo, Gen, Wang, Wenhai, Mu, Yao, Hou, Zhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
von: Chen, Junting, et al.
Veröffentlicht: (2025)
von: Chen, Junting, et al.
Veröffentlicht: (2025)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
von: Shen, Weijie, et al.
Veröffentlicht: (2025)
von: Shen, Weijie, et al.
Veröffentlicht: (2025)
ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework
von: Chen, Guanzhou, et al.
Veröffentlicht: (2026)
von: Chen, Guanzhou, et al.
Veröffentlicht: (2026)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
von: Lei, Zhenxin, et al.
Veröffentlicht: (2025)
von: Lei, Zhenxin, et al.
Veröffentlicht: (2025)
CoMemo: LVLMs Need Image Context with Image Memory
von: Liu, Shi, et al.
Veröffentlicht: (2025)
von: Liu, Shi, et al.
Veröffentlicht: (2025)
ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution
von: Cui, Long, et al.
Veröffentlicht: (2025)
von: Cui, Long, et al.
Veröffentlicht: (2025)
GenExam: A Multidisciplinary Text-to-Image Exam
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
von: Liu, Yangzhou, et al.
Veröffentlicht: (2024)
von: Liu, Yangzhou, et al.
Veröffentlicht: (2024)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
Docopilot: Improving Multimodal Models for Document-Level Understanding
von: Duan, Yuchen, et al.
Veröffentlicht: (2025)
von: Duan, Yuchen, et al.
Veröffentlicht: (2025)
HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System
von: Yang, Tianshuo, et al.
Veröffentlicht: (2026)
von: Yang, Tianshuo, et al.
Veröffentlicht: (2026)
Demystify Transformers & Convolutions in Modern Image Deep Networks
von: Hu, Xiaowei, et al.
Veröffentlicht: (2022)
von: Hu, Xiaowei, et al.
Veröffentlicht: (2022)
Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments
von: Yang, Yang, et al.
Veröffentlicht: (2024)
von: Yang, Yang, et al.
Veröffentlicht: (2024)
Sequential Diffusion Language Models
von: Liu, Yangzhou, et al.
Veröffentlicht: (2025)
von: Liu, Yangzhou, et al.
Veröffentlicht: (2025)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
von: Lv, Qi, et al.
Veröffentlicht: (2025)
von: Lv, Qi, et al.
Veröffentlicht: (2025)
Language-to-Space Programming for Training-Free 3D Visual Grounding
von: Mi, Boyu, et al.
Veröffentlicht: (2025)
von: Mi, Boyu, et al.
Veröffentlicht: (2025)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
von: Zhai, Shaopeng, et al.
Veröffentlicht: (2025)
von: Zhai, Shaopeng, et al.
Veröffentlicht: (2025)
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
von: Ren, Yiming, et al.
Veröffentlicht: (2025)
von: Ren, Yiming, et al.
Veröffentlicht: (2025)
P-RAG: Progressive Retrieval Augmented Generation For Planning on Embodied Everyday Task
von: Xu, Weiye, et al.
Veröffentlicht: (2024)
von: Xu, Weiye, et al.
Veröffentlicht: (2024)
EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence
von: Wang, Xinjie, et al.
Veröffentlicht: (2025)
von: Wang, Xinjie, et al.
Veröffentlicht: (2025)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit
von: Ben, Qingwei, et al.
Veröffentlicht: (2025)
von: Ben, Qingwei, et al.
Veröffentlicht: (2025)
Demystifying Action Space Design for Robotic Manipulation Policies
von: Feng, Yuchun, et al.
Veröffentlicht: (2026)
von: Feng, Yuchun, et al.
Veröffentlicht: (2026)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
Nimbus: A Unified Embodied Synthetic Data Generation Framework
von: He, Zeyu, et al.
Veröffentlicht: (2026)
von: He, Zeyu, et al.
Veröffentlicht: (2026)
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
von: Huang, Haifeng, et al.
Veröffentlicht: (2026)
von: Huang, Haifeng, et al.
Veröffentlicht: (2026)
GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction
von: Chen, Xiao, et al.
Veröffentlicht: (2024)
von: Chen, Xiao, et al.
Veröffentlicht: (2024)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024)
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024)
Low-Resolution Action Recognition for Tiny Actions Challenge
von: Chen, Boyu, et al.
Veröffentlicht: (2022)
von: Chen, Boyu, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
von: Luo, Gen, et al.
Veröffentlicht: (2025) -
OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
von: Chen, Junting, et al.
Veröffentlicht: (2025) -
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025) -
Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
von: Shen, Weijie, et al.
Veröffentlicht: (2025) -
ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework
von: Chen, Guanzhou, et al.
Veröffentlicht: (2026)