How Foundational Skills Influence VLM-based Embodied Agents:A Native Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Bo, Bu, Pi, Pan, Keyu, Xu, Xinrun, Zhao, Yinxiu, Chen, Miao, Du, Yang, Li, Lin, Song, Jun, Xu, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoJSON Agents:A Multi-Agent LLM Architecture for Geospatial Analysis-Function Calling vs Code Generation
by: Luo, Qianqian, et al.
Published: (2025)
by: Luo, Qianqian, et al.
Published: (2025)
MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
by: Yue, Junpeng, et al.
Published: (2024)
by: Yue, Junpeng, et al.
Published: (2024)
Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case
by: Chen, Peng, et al.
Published: (2024)
by: Chen, Peng, et al.
Published: (2024)
DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
by: Xu, Xinrun, et al.
Published: (2025)
by: Xu, Xinrun, et al.
Published: (2025)
The Agent Use of Agent Beings: Agent Cybernetics Is the Missing Science of Foundation Agents
by: Wang, Xinrun, et al.
Published: (2026)
by: Wang, Xinrun, et al.
Published: (2026)
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
by: Ju, Ruofei, et al.
Published: (2026)
by: Ju, Ruofei, et al.
Published: (2026)
ICPRL: Acquiring Physical Intuition from Interactive Control
by: Xu, Xinrun, et al.
Published: (2026)
by: Xu, Xinrun, et al.
Published: (2026)
InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
by: Ai, Qihang, et al.
Published: (2025)
by: Ai, Qihang, et al.
Published: (2025)
GEMS: Agent-Native Multimodal Generation with Memory and Skills
by: He, Zefeng, et al.
Published: (2026)
by: He, Zefeng, et al.
Published: (2026)
ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
by: Wang, Yichen, et al.
Published: (2025)
by: Wang, Yichen, et al.
Published: (2025)
EMAC+: Embodied Multimodal Agent for Collaborative Planning with VLM+LLM
by: Ao, Shuang, et al.
Published: (2025)
by: Ao, Shuang, et al.
Published: (2025)
An Atomic Skill Library Construction Method for Data-Efficient Embodied Manipulation
by: Li, Dongjiang, et al.
Published: (2025)
by: Li, Dongjiang, et al.
Published: (2025)
MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
by: Zhang, Ruoxuan, et al.
Published: (2025)
by: Zhang, Ruoxuan, et al.
Published: (2025)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
by: Zhan, Qiusi, et al.
Published: (2025)
by: Zhan, Qiusi, et al.
Published: (2025)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
by: Xie, Yiping, et al.
Published: (2026)
by: Xie, Yiping, et al.
Published: (2026)
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
by: Bai, Shuanghao, et al.
Published: (2025)
by: Bai, Shuanghao, et al.
Published: (2025)
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
by: Wang, Pan, et al.
Published: (2026)
by: Wang, Pan, et al.
Published: (2026)
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
by: Suglia, Alessandro, et al.
Published: (2024)
by: Suglia, Alessandro, et al.
Published: (2024)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
SELU: Self-Learning Embodied MLLMs in Unknown Environments
by: Li, Boyu, et al.
Published: (2024)
by: Li, Boyu, et al.
Published: (2024)
VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
by: Sarch, Gabriel, et al.
Published: (2024)
by: Sarch, Gabriel, et al.
Published: (2024)
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
by: Xu, Lu, et al.
Published: (2025)
by: Xu, Lu, et al.
Published: (2025)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
by: Ahn, Michael, et al.
Published: (2024)
by: Ahn, Michael, et al.
Published: (2024)
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents
by: X, Tencent Robotics, et al.
Published: (2026)
by: X, Tencent Robotics, et al.
Published: (2026)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
by: Tan, Weihao, et al.
Published: (2024)
by: Tan, Weihao, et al.
Published: (2024)
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
by: V Team, et al.
Published: (2026)
by: V Team, et al.
Published: (2026)
VLM-CAD: VLM-Optimized Collaborative Agent Design Workflow for Analog Circuit Sizing
by: Pan, Guanyuan, et al.
Published: (2026)
by: Pan, Guanyuan, et al.
Published: (2026)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
by: Lu, Xiaoya, et al.
Published: (2025)
by: Lu, Xiaoya, et al.
Published: (2025)
Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills
by: Yuan, Haoqi, et al.
Published: (2025)
by: Yuan, Haoqi, et al.
Published: (2025)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
by: Li, Boyu, et al.
Published: (2026)
by: Li, Boyu, et al.
Published: (2026)
Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
by: Sun, Nan, et al.
Published: (2025)
by: Sun, Nan, et al.
Published: (2025)
Agent Skill Framework: Perspectives on the Potential of Small Language Models in Industrial Environments
by: Xu, Yangjie, et al.
Published: (2026)
by: Xu, Yangjie, et al.
Published: (2026)
Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling
by: Liu, Gongye, et al.
Published: (2026)
by: Liu, Gongye, et al.
Published: (2026)
HumanVLM: Foundation for Human-Scene Vision-Language Model
by: Dai, Dawei, et al.
Published: (2024)
by: Dai, Dawei, et al.
Published: (2024)
Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own
by: Ye, Weirui, et al.
Published: (2023)
by: Ye, Weirui, et al.
Published: (2023)
Knows: Agent-Native Structured Research Representations
by: Yu, Guangsheng, et al.
Published: (2026)
by: Yu, Guangsheng, et al.
Published: (2026)
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
by: Dissanayake, Dinura, et al.
Published: (2025)
by: Dissanayake, Dinura, et al.
Published: (2025)
AgentStudio: A Toolkit for Building General Virtual Agents
by: Zheng, Longtao, et al.
Published: (2024)
by: Zheng, Longtao, et al.
Published: (2024)
SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces
by: Xu, Duling, et al.
Published: (2026)
by: Xu, Duling, et al.
Published: (2026)
Similar Items
-
GeoJSON Agents:A Multi-Agent LLM Architecture for Geospatial Analysis-Function Calling vs Code Generation
by: Luo, Qianqian, et al.
Published: (2025) -
MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
by: Yue, Junpeng, et al.
Published: (2024) -
Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case
by: Chen, Peng, et al.
Published: (2024) -
DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
by: Xu, Xinrun, et al.
Published: (2025) -
The Agent Use of Agent Beings: Agent Cybernetics Is the Missing Science of Foundation Agents
by: Wang, Xinrun, et al.
Published: (2026)