A Low-Cost Vision-Based Tactile Gripper with Pretraining Learning for Contact-Rich Manipulation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Yaohua, Ou, Binkai, Qiu, Zicheng, Hao, Ce, Zhang, Hengjun
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914303124176896
author Liu, Yaohua
Ou, Binkai
Qiu, Zicheng
Hao, Ce
Zhang, Hengjun
author_facet Liu, Yaohua
Ou, Binkai
Qiu, Zicheng
Hao, Ce
Zhang, Hengjun
contents Robotic manipulation in contact-rich environments remains challenging, particularly when relying on conventional tactile sensors that suffer from limited sensing range, reliability, and cost-effectiveness. In this work, we present LVTG, a low-cost visuo-tactile gripper designed for stable, robust, and efficient physical interaction. Unlike existing visuo-tactile sensors, LVTG enables more effective and stable grasping of larger and heavier everyday objects, thanks to its enhanced tactile sensing area and greater opening angle. Its surface skin is made of highly wear-resistant material, significantly improving durability and extending operational lifespan. The integration of vision and tactile feedback allows LVTG to provide rich, high-fidelity sensory data, facilitating reliable perception during complex manipulation tasks. Furthermore, LVTG features a modular design that supports rapid maintenance and replacement. To effectively fuse vision and touch, We adopt a CLIP-inspired contrastive learning objective to align tactile embeddings with their corresponding visual observations, enabling a shared cross-modal representation space for visuo-tactile perception. This alignment improves the performance of an Action Chunking Transformer (ACT) policy in contact-rich manipulation, leading to more efficient data collection and more effective policy learning. Compared to the original ACT method, the proposed LVTG with pretraining achieves significantly higher success rates in manipulation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00514
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Low-Cost Vision-Based Tactile Gripper with Pretraining Learning for Contact-Rich Manipulation
Liu, Yaohua
Ou, Binkai
Qiu, Zicheng
Hao, Ce
Zhang, Hengjun
Robotics
Robotic manipulation in contact-rich environments remains challenging, particularly when relying on conventional tactile sensors that suffer from limited sensing range, reliability, and cost-effectiveness. In this work, we present LVTG, a low-cost visuo-tactile gripper designed for stable, robust, and efficient physical interaction. Unlike existing visuo-tactile sensors, LVTG enables more effective and stable grasping of larger and heavier everyday objects, thanks to its enhanced tactile sensing area and greater opening angle. Its surface skin is made of highly wear-resistant material, significantly improving durability and extending operational lifespan. The integration of vision and tactile feedback allows LVTG to provide rich, high-fidelity sensory data, facilitating reliable perception during complex manipulation tasks. Furthermore, LVTG features a modular design that supports rapid maintenance and replacement. To effectively fuse vision and touch, We adopt a CLIP-inspired contrastive learning objective to align tactile embeddings with their corresponding visual observations, enabling a shared cross-modal representation space for visuo-tactile perception. This alignment improves the performance of an Action Chunking Transformer (ACT) policy in contact-rich manipulation, leading to more efficient data collection and more effective policy learning. Compared to the original ACT method, the proposed LVTG with pretraining achieves significantly higher success rates in manipulation tasks.
title A Low-Cost Vision-Based Tactile Gripper with Pretraining Learning for Contact-Rich Manipulation
topic Robotics
url https://arxiv.org/abs/2602.00514