EgoLM: Multi-Modal Language Model of Egocentric Motions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hong, Fangzhou, Guzov, Vladimir, Kim, Hyo Jin, Ye, Yuting, Newcombe, Richard, Liu, Ziwei, Ma, Lingni
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929516315672576
author Hong, Fangzhou
Guzov, Vladimir
Kim, Hyo Jin
Ye, Yuting
Newcombe, Richard
Liu, Ziwei
Ma, Lingni
author_facet Hong, Fangzhou
Guzov, Vladimir
Kim, Hyo Jin
Ye, Yuting
Newcombe, Richard
Liu, Ziwei
Ma, Lingni
contents As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and understands egocentric motions from multi-modal inputs, e.g., egocentric videos and motion sensors. EgoLM exploits rich contexts for the disambiguation of egomotion tracking and understanding, which are ill-posed under single modality conditions. To facilitate the versatile and multi-modal framework, our key insight is to model the joint distribution of egocentric motions and natural languages using large language models (LLM). Multi-modal sensor inputs are encoded and projected to the joint latent space of language models, and used to prompt motion generation or text generation for egomotion tracking or understanding, respectively. Extensive experiments on large-scale multi-modal human motion dataset validate the effectiveness of EgoLM as a generalist model for universal egocentric learning.
format Preprint
id arxiv_https___arxiv_org_abs_2409_18127
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EgoLM: Multi-Modal Language Model of Egocentric Motions
Hong, Fangzhou
Guzov, Vladimir
Kim, Hyo Jin
Ye, Yuting
Newcombe, Richard
Liu, Ziwei
Ma, Lingni
Computer Vision and Pattern Recognition
As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and understands egocentric motions from multi-modal inputs, e.g., egocentric videos and motion sensors. EgoLM exploits rich contexts for the disambiguation of egomotion tracking and understanding, which are ill-posed under single modality conditions. To facilitate the versatile and multi-modal framework, our key insight is to model the joint distribution of egocentric motions and natural languages using large language models (LLM). Multi-modal sensor inputs are encoded and projected to the joint latent space of language models, and used to prompt motion generation or text generation for egomotion tracking or understanding, respectively. Extensive experiments on large-scale multi-modal human motion dataset validate the effectiveness of EgoLM as a generalist model for universal egocentric learning.
title EgoLM: Multi-Modal Language Model of Egocentric Motions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.18127