InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Xinyi, Chen, Yilun, Fu, Yanwei, Gao, Ning, Jia, Jiaya, Jin, Weiyang, Li, Hao, Mu, Yao, Pang, Jiangmiao, Qiao, Yu, Tian, Yang, Wang, Bin, Wang, Bolun, Wang, Fangjing, Wang, Hanqing, Wang, Tai, Wang, Ziqin, Wei, Xueyuan, Wu, Chao, Yang, Shuai, Ye, Jinhui, Yu, Junqiu, Zeng, Jia, Zhang, Jingjing, Zhang, Jinyu, Zhang, Shi, Zheng, Feng, Zhou, Bowen, Zhu, Yangkun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
by: Cai, Junhao, et al.
Published: (2026)
by: Cai, Junhao, et al.
Published: (2026)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
Language-to-Space Programming for Training-Free 3D Visual Grounding
by: Mi, Boyu, et al.
Published: (2025)
by: Mi, Boyu, et al.
Published: (2025)
Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation
by: Han, Xiaoshen, et al.
Published: (2025)
by: Han, Xiaoshen, et al.
Published: (2025)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
OVExp: Open Vocabulary Exploration for Object-Oriented Navigation
by: Wei, Meng, et al.
Published: (2024)
by: Wei, Meng, et al.
Published: (2024)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance
by: Cai, Wenzhe, et al.
Published: (2025)
by: Cai, Wenzhe, et al.
Published: (2025)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
by: Wang, Liuyi, et al.
Published: (2025)
by: Wang, Liuyi, et al.
Published: (2025)
GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
by: Gao, Ning, et al.
Published: (2025)
by: Gao, Ning, et al.
Published: (2025)
InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
by: Tian, Yang, et al.
Published: (2025)
by: Tian, Yang, et al.
Published: (2025)
InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
by: Zhong, Weipeng, et al.
Published: (2025)
by: Zhong, Weipeng, et al.
Published: (2025)
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
by: Xu, Xiaoxu, et al.
Published: (2026)
by: Xu, Xiaoxu, et al.
Published: (2026)
PointLLM: Empowering Large Language Models to Understand Point Clouds
by: Xu, Runsen, et al.
Published: (2023)
by: Xu, Runsen, et al.
Published: (2023)
Dimension-reduced Optimization of Multi-zone Thermostatically Controlled Loads
by: Cui, Xueyuan, et al.
Published: (2025)
by: Cui, Xueyuan, et al.
Published: (2025)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
by: Li, Haozhan, et al.
Published: (2025)
by: Li, Haozhan, et al.
Published: (2025)
Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
by: Zheng, Jinliang, et al.
Published: (2025)
by: Zheng, Jinliang, et al.
Published: (2025)
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
by: Wei, Meng, et al.
Published: (2025)
by: Wei, Meng, et al.
Published: (2025)
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
by: Wang, Haomin, et al.
Published: (2025)
by: Wang, Haomin, et al.
Published: (2025)
RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
by: Li, Yanwei, et al.
Published: (2024)
by: Li, Yanwei, et al.
Published: (2024)
EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation
by: Cao, Jiajun, et al.
Published: (2026)
by: Cao, Jiajun, et al.
Published: (2026)
Disentangling nuclear structure through multiparticle azimuthal correlations in high-energy isobar collisions
by: Wang, Zaining, et al.
Published: (2024)
by: Wang, Zaining, et al.
Published: (2024)
Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
by: Wang, Yikai, et al.
Published: (2026)
by: Wang, Yikai, et al.
Published: (2026)
NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
by: Chen, Jiahong, et al.
Published: (2025)
by: Chen, Jiahong, et al.
Published: (2025)
Tac2Real: Reliable and GPU Visuotactile Simulation for Online Reinforcement Learning and Zero-Shot Real-World Deployment
by: Yan, Ningyu, et al.
Published: (2026)
by: Yan, Ningyu, et al.
Published: (2026)
MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
by: Fu, Yuxia, et al.
Published: (2025)
by: Fu, Yuxia, et al.
Published: (2025)
VisionZip: Longer is Better but Not Necessary in Vision Language Models
by: Yang, Senqiao, et al.
Published: (2024)
by: Yang, Senqiao, et al.
Published: (2024)
Nimbus: A Unified Embodied Synthetic Data Generation Framework
by: He, Zeyu, et al.
Published: (2026)
by: He, Zeyu, et al.
Published: (2026)
TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers
by: Yu, Bin, et al.
Published: (2026)
by: Yu, Bin, et al.
Published: (2026)
Local Differential Privacy for Distributed Stochastic Aggregative Optimization with Guaranteed Optimality
by: Chen, Ziqin, et al.
Published: (2025)
by: Chen, Ziqin, et al.
Published: (2025)
Gradient Manipulation in Distributed Stochastic Gradient Descent with Strategic Agents: Truthful Incentives with Convergence Guarantees
by: Chen, Ziqin, et al.
Published: (2026)
by: Chen, Ziqin, et al.
Published: (2026)
Privacy-Preserving Distributed Optimization and Learning
by: Chen, Ziqin, et al.
Published: (2024)
by: Chen, Ziqin, et al.
Published: (2024)
Locally Differentially Private Distributed Online Learning with Guaranteed Optimality
by: Chen, Ziqin, et al.
Published: (2023)
by: Chen, Ziqin, et al.
Published: (2023)
Grounded 3D-LLM with Referent Tokens
by: Chen, Yilun, et al.
Published: (2024)
by: Chen, Yilun, et al.
Published: (2024)
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization
by: Zhu, Mingkang, et al.
Published: (2025)
by: Zhu, Mingkang, et al.
Published: (2025)
Similar Items
-
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
by: Ye, Jinhui, et al.
Published: (2026) -
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
by: Cai, Junhao, et al.
Published: (2026) -
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025) -
Language-to-Space Programming for Training-Free 3D Visual Grounding
by: Mi, Boyu, et al.
Published: (2025) -
Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation
by: Han, Xiaoshen, et al.
Published: (2025)