MonSTeR: a Unified Model for Motion, Scene, Text Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Collorone, Luca, Gioia, Matteo, Pappa, Massimiliano, Leoni, Paolo, Ficarra, Giovanni, Litany, Or, Spinelli, Indro, Galasso, Fabio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911190932783104
author Collorone, Luca
Gioia, Matteo
Pappa, Massimiliano
Leoni, Paolo
Ficarra, Giovanni
Litany, Or
Spinelli, Indro
Galasso, Fabio
author_facet Collorone, Luca
Gioia, Matteo
Pappa, Massimiliano
Leoni, Paolo
Ficarra, Giovanni
Litany, Or
Spinelli, Indro
Galasso, Fabio
contents Intention drives human movement in complex environments, but such movement can only happen if the surrounding context supports it. Despite the intuitive nature of this mechanism, existing research has not yet provided tools to evaluate the alignment between skeletal movement (motion), intention (text), and the surrounding context (scene). In this work, we introduce MonSTeR, the first MOtioN-Scene-TExt Retrieval model. Inspired by the modeling of higher-order relations, MonSTeR constructs a unified latent space by leveraging unimodal and cross-modal representations. This allows MonSTeR to capture the intricate dependencies between modalities, enabling flexible but robust retrieval across various tasks. Our results show that MonSTeR outperforms trimodal models that rely solely on unimodal representations. Furthermore, we validate the alignment of our retrieval scores with human preferences through a dedicated user study. We demonstrate the versatility of MonSTeR's latent space on zero-shot in-Scene Object Placement and Motion Captioning. Code and pre-trained models are available at github.com/colloroneluca/MonSTeR.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03200
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MonSTeR: a Unified Model for Motion, Scene, Text Retrieval
Collorone, Luca
Gioia, Matteo
Pappa, Massimiliano
Leoni, Paolo
Ficarra, Giovanni
Litany, Or
Spinelli, Indro
Galasso, Fabio
Computer Vision and Pattern Recognition
Intention drives human movement in complex environments, but such movement can only happen if the surrounding context supports it. Despite the intuitive nature of this mechanism, existing research has not yet provided tools to evaluate the alignment between skeletal movement (motion), intention (text), and the surrounding context (scene). In this work, we introduce MonSTeR, the first MOtioN-Scene-TExt Retrieval model. Inspired by the modeling of higher-order relations, MonSTeR constructs a unified latent space by leveraging unimodal and cross-modal representations. This allows MonSTeR to capture the intricate dependencies between modalities, enabling flexible but robust retrieval across various tasks. Our results show that MonSTeR outperforms trimodal models that rely solely on unimodal representations. Furthermore, we validate the alignment of our retrieval scores with human preferences through a dedicated user study. We demonstrate the versatility of MonSTeR's latent space on zero-shot in-Scene Object Placement and Motion Captioning. Code and pre-trained models are available at github.com/colloroneluca/MonSTeR.
title MonSTeR: a Unified Model for Motion, Scene, Text Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.03200