Dense Motion Captioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Shiyao, Liberatori, Benedetta, Varol, Gül, Rota, Paolo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918190791000064
author Xu, Shiyao
Liberatori, Benedetta
Varol, Gül
Rota, Paolo
author_facet Xu, Shiyao
Liberatori, Benedetta
Varol, Gül
Rota, Paolo
contents Recent advances in 3D human motion and language integration have primarily focused on text-to-motion generation, leaving the task of motion understanding relatively unexplored. We introduce Dense Motion Captioning, a novel task that aims to temporally localize and caption actions within 3D human motion sequences. Current datasets fall short in providing detailed temporal annotations and predominantly consist of short sequences featuring few actions. To overcome these limitations, we present the Complex Motion Dataset (CompMo), the first large-scale dataset featuring richly annotated, complex motion sequences with precise temporal boundaries. Built through a carefully designed data generation pipeline, CompMo includes 60,000 motion sequences, each composed of multiple actions ranging from at least two to ten, accurately annotated with their temporal extents. We further present DEMO, a model that integrates a large language model with a simple motion adapter, trained to generate dense, temporally grounded captions. Our experiments show that DEMO substantially outperforms existing methods on CompMo as well as on adapted benchmarks, establishing a robust baseline for future research in 3D motion understanding and captioning.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05369
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dense Motion Captioning
Xu, Shiyao
Liberatori, Benedetta
Varol, Gül
Rota, Paolo
Computer Vision and Pattern Recognition
I.2.10; I.4.8; I.5.4
Recent advances in 3D human motion and language integration have primarily focused on text-to-motion generation, leaving the task of motion understanding relatively unexplored. We introduce Dense Motion Captioning, a novel task that aims to temporally localize and caption actions within 3D human motion sequences. Current datasets fall short in providing detailed temporal annotations and predominantly consist of short sequences featuring few actions. To overcome these limitations, we present the Complex Motion Dataset (CompMo), the first large-scale dataset featuring richly annotated, complex motion sequences with precise temporal boundaries. Built through a carefully designed data generation pipeline, CompMo includes 60,000 motion sequences, each composed of multiple actions ranging from at least two to ten, accurately annotated with their temporal extents. We further present DEMO, a model that integrates a large language model with a simple motion adapter, trained to generate dense, temporally grounded captions. Our experiments show that DEMO substantially outperforms existing methods on CompMo as well as on adapted benchmarks, establishing a robust baseline for future research in 3D motion understanding and captioning.
title Dense Motion Captioning
topic Computer Vision and Pattern Recognition
I.2.10; I.4.8; I.5.4
url https://arxiv.org/abs/2511.05369