LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909826454388736 |
|---|---|
| author | Wang, Yihao Memmesheimer, Raphael Behnke, Sven |
| author_facet | Wang, Yihao Memmesheimer, Raphael Behnke, Sven |
| contents | The availability of large language models and open-vocabulary object perception methods enables more flexibility for domestic service robots. The large variability of domestic tasks can be addressed without implementing each task individually by providing the robot with a task description along with appropriate environment information. In this work, we propose LIAM - an end-to-end model that predicts action transcripts based on language, image, action, and map inputs. Language and image inputs are encoded with a CLIP backbone, for which we designed two pre-training tasks to fine-tune its weights and pre-align the latent spaces. We evaluate our method on the ALFRED dataset, a simulator-generated benchmark for domestic tasks. Our results demonstrate the importance of pre-aligning embedding spaces from different modalities and the efficacy of incorporating semantic maps. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_12230 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps Wang, Yihao Memmesheimer, Raphael Behnke, Sven Computer Vision and Pattern Recognition Artificial Intelligence Robotics The availability of large language models and open-vocabulary object perception methods enables more flexibility for domestic service robots. The large variability of domestic tasks can be addressed without implementing each task individually by providing the robot with a task description along with appropriate environment information. In this work, we propose LIAM - an end-to-end model that predicts action transcripts based on language, image, action, and map inputs. Language and image inputs are encoded with a CLIP backbone, for which we designed two pre-training tasks to fine-tune its weights and pre-align the latent spaces. We evaluate our method on the ALFRED dataset, a simulator-generated benchmark for domestic tasks. Our results demonstrate the importance of pre-aligning embedding spaces from different modalities and the efficacy of incorporating semantic maps. |
| title | LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Robotics |
| url | https://arxiv.org/abs/2503.12230 |