VLM-driven Skill Selection for Robotic Assembly Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Jeong-Jung, Koh, Doo-Yeol, Kim, Chang-Hyun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909893991071744
author Kim, Jeong-Jung
Koh, Doo-Yeol
Kim, Chang-Hyun
author_facet Kim, Jeong-Jung
Koh, Doo-Yeol
Kim, Chang-Hyun
contents This paper presents a robotic assembly framework that combines Vision-Language Models (VLMs) with imitation learning for assembly manipulation tasks. Our system employs a gripper-equipped robot that moves in 3D space to perform assembly operations. The framework integrates visual perception, natural language understanding, and learned primitive skills to enable flexible and adaptive robotic manipulation. Experimental results demonstrate the effectiveness of our approach in assembly scenarios, achieving high success rates while maintaining interpretability through the structured primitive skill decomposition.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05680
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLM-driven Skill Selection for Robotic Assembly Tasks
Kim, Jeong-Jung
Koh, Doo-Yeol
Kim, Chang-Hyun
Robotics
This paper presents a robotic assembly framework that combines Vision-Language Models (VLMs) with imitation learning for assembly manipulation tasks. Our system employs a gripper-equipped robot that moves in 3D space to perform assembly operations. The framework integrates visual perception, natural language understanding, and learned primitive skills to enable flexible and adaptive robotic manipulation. Experimental results demonstrate the effectiveness of our approach in assembly scenarios, achieving high success rates while maintaining interpretability through the structured primitive skill decomposition.
title VLM-driven Skill Selection for Robotic Assembly Tasks
topic Robotics
url https://arxiv.org/abs/2511.05680