FTACT: Force Torque aware Action Chunking Transformer for Pick-and-Reorient Bottle Task

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Watanabe, Ryo, Alvarez, Maxime, Ferreiro, Pablo, Savkin, Pavel, Sano, Genki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915519059197952
author Watanabe, Ryo
Alvarez, Maxime
Ferreiro, Pablo
Savkin, Pavel
Sano, Genki
author_facet Watanabe, Ryo
Alvarez, Maxime
Ferreiro, Pablo
Savkin, Pavel
Sano, Genki
contents Manipulator robots are increasingly being deployed in retail environments, yet contact rich edge cases still trigger costly human teleoperation. A prominent example is upright lying beverage bottles, where purely visual cues are often insufficient to resolve subtle contact events required for precise manipulation. We present a multimodal Imitation Learning policy that augments the Action Chunking Transformer with force and torque sensing, enabling end-to-end learning over images, joint states, and forces and torques. Deployed on Ghost, single-arm platform by Telexistence Inc, our approach improves Pick-and-Reorient bottle task by detecting and exploiting contact transitions during pressing and placement. Hardware experiments demonstrate greater task success compared to baseline matching the observation space of ACT as an ablation and experiments indicate that force and torque signals are beneficial in the press and place phases where visual observability is limited, supporting the use of interaction forces as a complementary modality for contact rich skills. The results suggest a practical path to scaling retail manipulation by combining modern imitation learning architectures with lightweight force and torque sensing.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23112
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FTACT: Force Torque aware Action Chunking Transformer for Pick-and-Reorient Bottle Task
Watanabe, Ryo
Alvarez, Maxime
Ferreiro, Pablo
Savkin, Pavel
Sano, Genki
Robotics
Manipulator robots are increasingly being deployed in retail environments, yet contact rich edge cases still trigger costly human teleoperation. A prominent example is upright lying beverage bottles, where purely visual cues are often insufficient to resolve subtle contact events required for precise manipulation. We present a multimodal Imitation Learning policy that augments the Action Chunking Transformer with force and torque sensing, enabling end-to-end learning over images, joint states, and forces and torques. Deployed on Ghost, single-arm platform by Telexistence Inc, our approach improves Pick-and-Reorient bottle task by detecting and exploiting contact transitions during pressing and placement. Hardware experiments demonstrate greater task success compared to baseline matching the observation space of ACT as an ablation and experiments indicate that force and torque signals are beneficial in the press and place phases where visual observability is limited, supporting the use of interaction forces as a complementary modality for contact rich skills. The results suggest a practical path to scaling retail manipulation by combining modern imitation learning architectures with lightweight force and torque sensing.
title FTACT: Force Torque aware Action Chunking Transformer for Pick-and-Reorient Bottle Task
topic Robotics
url https://arxiv.org/abs/2509.23112