A Multi-Modal Interaction Framework for Efficient Human-Robot Collaborative Shelf Picking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pathak, Abhinav, Venkatesan, Kalaichelvi, Taha, Tarek, Muthusamy, Rajkumar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917981194289152
author Pathak, Abhinav
Venkatesan, Kalaichelvi
Taha, Tarek
Muthusamy, Rajkumar
author_facet Pathak, Abhinav
Venkatesan, Kalaichelvi
Taha, Tarek
Muthusamy, Rajkumar
contents The growing presence of service robots in human-centric environments, such as warehouses, demands seamless and intuitive human-robot collaboration. In this paper, we propose a collaborative shelf-picking framework that combines multimodal interaction, physics-based reasoning, and task division for enhanced human-robot teamwork. The framework enables the robot to recognize human pointing gestures, interpret verbal cues and voice commands, and communicate through visual and auditory feedback. Moreover, it is powered by a Large Language Model (LLM) which utilizes Chain of Thought (CoT) and a physics-based simulation engine for safely retrieving cluttered stacks of boxes on shelves, relationship graph for sub-task generation, extraction sequence planning and decision making. Furthermore, we validate the framework through real-world shelf picking experiments such as 1) Gesture-Guided Box Extraction, 2) Collaborative Shelf Clearing and 3) Collaborative Stability Assistance.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06593
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Multi-Modal Interaction Framework for Efficient Human-Robot Collaborative Shelf Picking
Pathak, Abhinav
Venkatesan, Kalaichelvi
Taha, Tarek
Muthusamy, Rajkumar
Robotics
Human-Computer Interaction
The growing presence of service robots in human-centric environments, such as warehouses, demands seamless and intuitive human-robot collaboration. In this paper, we propose a collaborative shelf-picking framework that combines multimodal interaction, physics-based reasoning, and task division for enhanced human-robot teamwork. The framework enables the robot to recognize human pointing gestures, interpret verbal cues and voice commands, and communicate through visual and auditory feedback. Moreover, it is powered by a Large Language Model (LLM) which utilizes Chain of Thought (CoT) and a physics-based simulation engine for safely retrieving cluttered stacks of boxes on shelves, relationship graph for sub-task generation, extraction sequence planning and decision making. Furthermore, we validate the framework through real-world shelf picking experiments such as 1) Gesture-Guided Box Extraction, 2) Collaborative Shelf Clearing and 3) Collaborative Stability Assistance.
title A Multi-Modal Interaction Framework for Efficient Human-Robot Collaborative Shelf Picking
topic Robotics
Human-Computer Interaction
url https://arxiv.org/abs/2504.06593