How Foundational Skills Influence VLM-based Embodied Agents:A Native Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | Peng, Bo, Bu, Pi, Pan, Keyu, Xu, Xinrun, Zhao, Yinxiu, Chen, Miao, Du, Yang, Li, Lin, Song, Jun, Xu, Tong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GeoJSON Agents:A Multi-Agent LLM Architecture for Geospatial Analysis-Function Calling vs Code Generation
di: Luo, Qianqian, et al.
Pubblicazione: (2025)
di: Luo, Qianqian, et al.
Pubblicazione: (2025)
MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
di: Yue, Junpeng, et al.
Pubblicazione: (2024)
di: Yue, Junpeng, et al.
Pubblicazione: (2024)
Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case
di: Chen, Peng, et al.
Pubblicazione: (2024)
di: Chen, Peng, et al.
Pubblicazione: (2024)
DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
di: Xu, Xinrun, et al.
Pubblicazione: (2025)
di: Xu, Xinrun, et al.
Pubblicazione: (2025)
The Agent Use of Agent Beings: Agent Cybernetics Is the Missing Science of Foundation Agents
di: Wang, Xinrun, et al.
Pubblicazione: (2026)
di: Wang, Xinrun, et al.
Pubblicazione: (2026)
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
di: Gu, Jihao, et al.
Pubblicazione: (2025)
di: Gu, Jihao, et al.
Pubblicazione: (2025)
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
di: Ju, Ruofei, et al.
Pubblicazione: (2026)
di: Ju, Ruofei, et al.
Pubblicazione: (2026)
ICPRL: Acquiring Physical Intuition from Interactive Control
di: Xu, Xinrun, et al.
Pubblicazione: (2026)
di: Xu, Xinrun, et al.
Pubblicazione: (2026)
InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
di: Ai, Qihang, et al.
Pubblicazione: (2025)
di: Ai, Qihang, et al.
Pubblicazione: (2025)
GEMS: Agent-Native Multimodal Generation with Memory and Skills
di: He, Zefeng, et al.
Pubblicazione: (2026)
di: He, Zefeng, et al.
Pubblicazione: (2026)
ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
di: Wang, Yichen, et al.
Pubblicazione: (2025)
di: Wang, Yichen, et al.
Pubblicazione: (2025)
EMAC+: Embodied Multimodal Agent for Collaborative Planning with VLM+LLM
di: Ao, Shuang, et al.
Pubblicazione: (2025)
di: Ao, Shuang, et al.
Pubblicazione: (2025)
An Atomic Skill Library Construction Method for Data-Efficient Embodied Manipulation
di: Li, Dongjiang, et al.
Pubblicazione: (2025)
di: Li, Dongjiang, et al.
Pubblicazione: (2025)
MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
di: Zhang, Ruoxuan, et al.
Pubblicazione: (2025)
di: Zhang, Ruoxuan, et al.
Pubblicazione: (2025)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
di: Zhan, Qiusi, et al.
Pubblicazione: (2025)
di: Zhan, Qiusi, et al.
Pubblicazione: (2025)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
di: Xie, Yiping, et al.
Pubblicazione: (2026)
di: Xie, Yiping, et al.
Pubblicazione: (2026)
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
di: Wang, Pan, et al.
Pubblicazione: (2026)
di: Wang, Pan, et al.
Pubblicazione: (2026)
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
di: Suglia, Alessandro, et al.
Pubblicazione: (2024)
di: Suglia, Alessandro, et al.
Pubblicazione: (2024)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
di: Xu, Runsen, et al.
Pubblicazione: (2024)
di: Xu, Runsen, et al.
Pubblicazione: (2024)
SELU: Self-Learning Embodied MLLMs in Unknown Environments
di: Li, Boyu, et al.
Pubblicazione: (2024)
di: Li, Boyu, et al.
Pubblicazione: (2024)
VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
di: Sarch, Gabriel, et al.
Pubblicazione: (2024)
di: Sarch, Gabriel, et al.
Pubblicazione: (2024)
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
di: Xu, Lu, et al.
Pubblicazione: (2025)
di: Xu, Lu, et al.
Pubblicazione: (2025)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
di: Ahn, Michael, et al.
Pubblicazione: (2024)
di: Ahn, Michael, et al.
Pubblicazione: (2024)
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents
di: X, Tencent Robotics, et al.
Pubblicazione: (2026)
di: X, Tencent Robotics, et al.
Pubblicazione: (2026)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
di: Tan, Weihao, et al.
Pubblicazione: (2024)
di: Tan, Weihao, et al.
Pubblicazione: (2024)
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
di: V Team, et al.
Pubblicazione: (2026)
di: V Team, et al.
Pubblicazione: (2026)
VLM-CAD: VLM-Optimized Collaborative Agent Design Workflow for Analog Circuit Sizing
di: Pan, Guanyuan, et al.
Pubblicazione: (2026)
di: Pan, Guanyuan, et al.
Pubblicazione: (2026)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
di: Lu, Xiaoya, et al.
Pubblicazione: (2025)
di: Lu, Xiaoya, et al.
Pubblicazione: (2025)
Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills
di: Yuan, Haoqi, et al.
Pubblicazione: (2025)
di: Yuan, Haoqi, et al.
Pubblicazione: (2025)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
di: Li, Boyu, et al.
Pubblicazione: (2026)
di: Li, Boyu, et al.
Pubblicazione: (2026)
Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
di: Sun, Nan, et al.
Pubblicazione: (2025)
di: Sun, Nan, et al.
Pubblicazione: (2025)
Agent Skill Framework: Perspectives on the Potential of Small Language Models in Industrial Environments
di: Xu, Yangjie, et al.
Pubblicazione: (2026)
di: Xu, Yangjie, et al.
Pubblicazione: (2026)
Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling
di: Liu, Gongye, et al.
Pubblicazione: (2026)
di: Liu, Gongye, et al.
Pubblicazione: (2026)
HumanVLM: Foundation for Human-Scene Vision-Language Model
di: Dai, Dawei, et al.
Pubblicazione: (2024)
di: Dai, Dawei, et al.
Pubblicazione: (2024)
Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own
di: Ye, Weirui, et al.
Pubblicazione: (2023)
di: Ye, Weirui, et al.
Pubblicazione: (2023)
Knows: Agent-Native Structured Research Representations
di: Yu, Guangsheng, et al.
Pubblicazione: (2026)
di: Yu, Guangsheng, et al.
Pubblicazione: (2026)
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
di: Dissanayake, Dinura, et al.
Pubblicazione: (2025)
di: Dissanayake, Dinura, et al.
Pubblicazione: (2025)
AgentStudio: A Toolkit for Building General Virtual Agents
di: Zheng, Longtao, et al.
Pubblicazione: (2024)
di: Zheng, Longtao, et al.
Pubblicazione: (2024)
SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces
di: Xu, Duling, et al.
Pubblicazione: (2026)
di: Xu, Duling, et al.
Pubblicazione: (2026)
Documenti analoghi
-
GeoJSON Agents:A Multi-Agent LLM Architecture for Geospatial Analysis-Function Calling vs Code Generation
di: Luo, Qianqian, et al.
Pubblicazione: (2025) -
MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
di: Yue, Junpeng, et al.
Pubblicazione: (2024) -
Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case
di: Chen, Peng, et al.
Pubblicazione: (2024) -
DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
di: Xu, Xinrun, et al.
Pubblicazione: (2025) -
The Agent Use of Agent Beings: Agent Cybernetics Is the Missing Science of Foundation Agents
di: Wang, Xinrun, et al.
Pubblicazione: (2026)