Saved in:
Bibliographic Details
Main Authors: Konstantinidou, Eleni, Kounalakis, Nikolaos, Efstathopoulos, Nikolaos, Papageorgiou, Dimitrios
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.07745
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913936311320576
author Konstantinidou, Eleni
Kounalakis, Nikolaos
Efstathopoulos, Nikolaos
Papageorgiou, Dimitrios
author_facet Konstantinidou, Eleni
Kounalakis, Nikolaos
Efstathopoulos, Nikolaos
Papageorgiou, Dimitrios
contents Despite their recent introduction to human society, Large Language Models (LLMs) have significantly affected the way we tackle mental challenges in our everyday lives. From optimizing our linguistic communication to assisting us in making important decisions, LLMs, such as ChatGPT, are notably reducing our cognitive load by gradually taking on an increasing share of our mental activities. In the context of Learning by Demonstration (LbD), classifying and segmenting complex motions into primitive actions, such as pushing, pulling, twisting etc, is considered to be a key-step towards encoding a task. In this work, we investigate the capabilities of LLMs to undertake this task, considering a finite set of predefined primitive actions found in fruit picking operations. By utilizing LLMs instead of simple supervised learning or analytic methods, we aim at making the method easily applicable and deployable in a real-life scenario. Three different fine-tuning approaches are investigated, compared on datasets captured kinesthetically, using a UR10e robot, during a fruit-picking scenario.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07745
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the capabilities of LLMs for classifying and segmenting time series of fruit picking motions into primitive actions
Konstantinidou, Eleni
Kounalakis, Nikolaos
Efstathopoulos, Nikolaos
Papageorgiou, Dimitrios
Robotics
Despite their recent introduction to human society, Large Language Models (LLMs) have significantly affected the way we tackle mental challenges in our everyday lives. From optimizing our linguistic communication to assisting us in making important decisions, LLMs, such as ChatGPT, are notably reducing our cognitive load by gradually taking on an increasing share of our mental activities. In the context of Learning by Demonstration (LbD), classifying and segmenting complex motions into primitive actions, such as pushing, pulling, twisting etc, is considered to be a key-step towards encoding a task. In this work, we investigate the capabilities of LLMs to undertake this task, considering a finite set of predefined primitive actions found in fruit picking operations. By utilizing LLMs instead of simple supervised learning or analytic methods, we aim at making the method easily applicable and deployable in a real-life scenario. Three different fine-tuning approaches are investigated, compared on datasets captured kinesthetically, using a UR10e robot, during a fruit-picking scenario.
title On the capabilities of LLMs for classifying and segmenting time series of fruit picking motions into primitive actions
topic Robotics
url https://arxiv.org/abs/2507.07745