A Touch, Vision, and Language Dataset for Multimodal Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Letian, Datta, Gaurav, Huang, Huang, Panitch, William Chung-Ho, Drake, Jaimyn, Ortiz, Joseph, Mukadam, Mustafa, Lambeta, Mike, Calandra, Roberto, Goldberg, Ken
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910337730609152
author Fu, Letian
Datta, Gaurav
Huang, Huang
Panitch, William Chung-Ho
Drake, Jaimyn
Ortiz, Joseph
Mukadam, Mustafa
Lambeta, Mike
Calandra, Roberto
Goldberg, Ken
author_facet Fu, Letian
Datta, Gaurav
Huang, Huang
Panitch, William Chung-Ho
Drake, Jaimyn
Ortiz, Joseph
Mukadam, Mustafa
Lambeta, Mike
Calandra, Roberto
Goldberg, Ken
contents Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model. This is partially due to the difficulty of obtaining natural language labels for tactile data and the complexity of aligning tactile readings with both visual observations and language descriptions. As a step towards bridging that gap, this work introduces a new dataset of 44K in-the-wild vision-touch pairs, with English language labels annotated by humans (10%) and textual pseudo-labels from GPT-4V (90%). We use this dataset to train a vision-language-aligned tactile encoder for open-vocabulary classification and a touch-vision-language (TVL) model for text generation using the trained encoder. Results suggest that by incorporating touch, the TVL model improves (+29% classification accuracy) touch-vision-language alignment over existing models trained on any pair of those modalities. Although only a small fraction of the dataset is human-labeled, the TVL model demonstrates improved visual-tactile understanding over GPT-4V (+12%) and open-source vision-language models (+32%) on a new touch-vision understanding benchmark. Code and data: https://tactile-vlm.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2402_13232
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Touch, Vision, and Language Dataset for Multimodal Alignment
Fu, Letian
Datta, Gaurav
Huang, Huang
Panitch, William Chung-Ho
Drake, Jaimyn
Ortiz, Joseph
Mukadam, Mustafa
Lambeta, Mike
Calandra, Roberto
Goldberg, Ken
Computer Vision and Pattern Recognition
Robotics
Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model. This is partially due to the difficulty of obtaining natural language labels for tactile data and the complexity of aligning tactile readings with both visual observations and language descriptions. As a step towards bridging that gap, this work introduces a new dataset of 44K in-the-wild vision-touch pairs, with English language labels annotated by humans (10%) and textual pseudo-labels from GPT-4V (90%). We use this dataset to train a vision-language-aligned tactile encoder for open-vocabulary classification and a touch-vision-language (TVL) model for text generation using the trained encoder. Results suggest that by incorporating touch, the TVL model improves (+29% classification accuracy) touch-vision-language alignment over existing models trained on any pair of those modalities. Although only a small fraction of the dataset is human-labeled, the TVL model demonstrates improved visual-tactile understanding over GPT-4V (+12%) and open-source vision-language models (+32%) on a new touch-vision understanding benchmark. Code and data: https://tactile-vlm.github.io.
title A Touch, Vision, and Language Dataset for Multimodal Alignment
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2402.13232