LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yihao, Memmesheimer, Raphael, Behnke, Sven
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909826454388736
author Wang, Yihao
Memmesheimer, Raphael
Behnke, Sven
author_facet Wang, Yihao
Memmesheimer, Raphael
Behnke, Sven
contents The availability of large language models and open-vocabulary object perception methods enables more flexibility for domestic service robots. The large variability of domestic tasks can be addressed without implementing each task individually by providing the robot with a task description along with appropriate environment information. In this work, we propose LIAM - an end-to-end model that predicts action transcripts based on language, image, action, and map inputs. Language and image inputs are encoded with a CLIP backbone, for which we designed two pre-training tasks to fine-tune its weights and pre-align the latent spaces. We evaluate our method on the ALFRED dataset, a simulator-generated benchmark for domestic tasks. Our results demonstrate the importance of pre-aligning embedding spaces from different modalities and the efficacy of incorporating semantic maps.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12230
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
Wang, Yihao
Memmesheimer, Raphael
Behnke, Sven
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
The availability of large language models and open-vocabulary object perception methods enables more flexibility for domestic service robots. The large variability of domestic tasks can be addressed without implementing each task individually by providing the robot with a task description along with appropriate environment information. In this work, we propose LIAM - an end-to-end model that predicts action transcripts based on language, image, action, and map inputs. Language and image inputs are encoded with a CLIP backbone, for which we designed two pre-training tasks to fine-tune its weights and pre-align the latent spaces. We evaluate our method on the ALFRED dataset, a simulator-generated benchmark for domestic tasks. Our results demonstrate the importance of pre-aligning embedding spaces from different modalities and the efficacy of incorporating semantic maps.
title LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2503.12230