MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Knauer, Markus, Fiorini, Edoardo, Mühlbauer, Maximilian, Schneyer, Stefan, Angsuratanawech, Promwat, Lay, Florian Samuel, Bachmann, Timo, Bustamante, Samuel, Nottensteiner, Korbinian, Stulp, Freek, Albu-Schäffer, Alin, Silvério, João, Eiband, Thomas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917430956130304
author Knauer, Markus
Fiorini, Edoardo
Mühlbauer, Maximilian
Schneyer, Stefan
Angsuratanawech, Promwat
Lay, Florian Samuel
Bachmann, Timo
Bustamante, Samuel
Nottensteiner, Korbinian
Stulp, Freek
Albu-Schäffer, Alin
Silvério, João
Eiband, Thomas
author_facet Knauer, Markus
Fiorini, Edoardo
Mühlbauer, Maximilian
Schneyer, Stefan
Angsuratanawech, Promwat
Lay, Florian Samuel
Bachmann, Timo
Bustamante, Samuel
Nottensteiner, Korbinian
Stulp, Freek
Albu-Schäffer, Alin
Silvério, João
Eiband, Thomas
contents Industrial robot applications require increasingly flexible systems that non-expert users can easily adapt for varying tasks and environments. However, different adaptations benefit from different interaction modalities. We present an interactive framework that enables robot skill adaptation through three complementary modalities: kinesthetic touch for precise spatial corrections, natural language for high-level semantic modifications, and a graphical web interface for visualizing geometric relations and trajectories, inspecting and adjusting parameters, and editing via-points by drag-and-drop. The framework integrates five components: energy-based human-intention detection, a tool-based LLM architecture (where the LLM selects and parameterizes predefined functions rather than generating code) for safe natural language adaptation, Kernelized Movement Primitives (KMPs) for motion encoding, probabilistic Virtual Fixtures for guided demonstration recording, and ergodic control for surface finishing. We demonstrate that this tool-based LLM architecture generalizes skill adaptation from KMPs to ergodic control, enabling voice-commanded surface finishing. Validation on a 7-DoF torque-controlled robot at the Automatica 2025 trade fair demonstrates the practical applicability of our approach in industrial settings.
format Preprint
id arxiv_https___arxiv_org_abs_2604_20468
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation
Knauer, Markus
Fiorini, Edoardo
Mühlbauer, Maximilian
Schneyer, Stefan
Angsuratanawech, Promwat
Lay, Florian Samuel
Bachmann, Timo
Bustamante, Samuel
Nottensteiner, Korbinian
Stulp, Freek
Albu-Schäffer, Alin
Silvério, João
Eiband, Thomas
Robotics
Artificial Intelligence
Computation and Language
Human-Computer Interaction
Machine Learning
Industrial robot applications require increasingly flexible systems that non-expert users can easily adapt for varying tasks and environments. However, different adaptations benefit from different interaction modalities. We present an interactive framework that enables robot skill adaptation through three complementary modalities: kinesthetic touch for precise spatial corrections, natural language for high-level semantic modifications, and a graphical web interface for visualizing geometric relations and trajectories, inspecting and adjusting parameters, and editing via-points by drag-and-drop. The framework integrates five components: energy-based human-intention detection, a tool-based LLM architecture (where the LLM selects and parameterizes predefined functions rather than generating code) for safe natural language adaptation, Kernelized Movement Primitives (KMPs) for motion encoding, probabilistic Virtual Fixtures for guided demonstration recording, and ergodic control for surface finishing. We demonstrate that this tool-based LLM architecture generalizes skill adaptation from KMPs to ergodic control, enabling voice-commanded surface finishing. Validation on a 7-DoF torque-controlled robot at the Automatica 2025 trade fair demonstrates the practical applicability of our approach in industrial settings.
title MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation
topic Robotics
Artificial Intelligence
Computation and Language
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2604.20468