Multimodal Visual-Tactile Representation Learning through Self-Supervised Contrastive Pre-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dave, Vedant, Lygerakis, Fotios, Rueckert, Elmar
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909079295754240
author Dave, Vedant
Lygerakis, Fotios
Rueckert, Elmar
author_facet Dave, Vedant
Lygerakis, Fotios
Rueckert, Elmar
contents The rapidly evolving field of robotics necessitates methods that can facilitate the fusion of multiple modalities. Specifically, when it comes to interacting with tangible objects, effectively combining visual and tactile sensory data is key to understanding and navigating the complex dynamics of the physical world, enabling a more nuanced and adaptable response to changing environments. Nevertheless, much of the earlier work in merging these two sensory modalities has relied on supervised methods utilizing datasets labeled by humans.This paper introduces MViTac, a novel methodology that leverages contrastive learning to integrate vision and touch sensations in a self-supervised fashion. By availing both sensory inputs, MViTac leverages intra and inter-modality losses for learning representations, resulting in enhanced material property classification and more adept grasping prediction. Through a series of experiments, we showcase the effectiveness of our method and its superiority over existing state-of-the-art self-supervised and supervised techniques. In evaluating our methodology, we focus on two distinct tasks: material classification and grasping success prediction. Our results indicate that MViTac facilitates the development of improved modality encoders, yielding more robust representations as evidenced by linear probing assessments.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12024
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal Visual-Tactile Representation Learning through Self-Supervised Contrastive Pre-Training
Dave, Vedant
Lygerakis, Fotios
Rueckert, Elmar
Robotics
Artificial Intelligence
Machine Learning
The rapidly evolving field of robotics necessitates methods that can facilitate the fusion of multiple modalities. Specifically, when it comes to interacting with tangible objects, effectively combining visual and tactile sensory data is key to understanding and navigating the complex dynamics of the physical world, enabling a more nuanced and adaptable response to changing environments. Nevertheless, much of the earlier work in merging these two sensory modalities has relied on supervised methods utilizing datasets labeled by humans.This paper introduces MViTac, a novel methodology that leverages contrastive learning to integrate vision and touch sensations in a self-supervised fashion. By availing both sensory inputs, MViTac leverages intra and inter-modality losses for learning representations, resulting in enhanced material property classification and more adept grasping prediction. Through a series of experiments, we showcase the effectiveness of our method and its superiority over existing state-of-the-art self-supervised and supervised techniques. In evaluating our methodology, we focus on two distinct tasks: material classification and grasping success prediction. Our results indicate that MViTac facilitates the development of improved modality encoders, yielding more robust representations as evidenced by linear probing assessments.
title Multimodal Visual-Tactile Representation Learning through Self-Supervised Contrastive Pre-Training
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2401.12024