SimpliHuMoN: Simplifying Human Motion Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agrawal, Aadya, Schwing, Alexander
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914368819560448
author Agrawal, Aadya
Schwing, Alexander
author_facet Agrawal, Aadya
Schwing, Alexander
contents Human motion prediction combines the tasks of trajectory forecasting and human pose prediction. For each of the two tasks, specialized models have been developed. Combining these models for holistic human motion prediction is non-trivial, and recent methods have struggled to compete on established benchmarks for individual tasks. To address this, we propose a simple yet effective transformer-based model for human motion prediction. The model employs a stack of self-attention modules to effectively capture both spatial dependencies within a pose and temporal relationships across a motion sequence. This simple, streamlined, end-to-end model is sufficiently versatile to handle pose-only, trajectory-only, and combined prediction tasks without task-specific modifications. We demonstrate that this approach achieves state-of-the-art results across all tasks through extensive experiments on a wide range of benchmark datasets, including Human3.6M, AMASS, ETH-UCY, and 3DPW.
format Preprint
id arxiv_https___arxiv_org_abs_2603_04399
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SimpliHuMoN: Simplifying Human Motion Prediction
Agrawal, Aadya
Schwing, Alexander
Computer Vision and Pattern Recognition
Machine Learning
Human motion prediction combines the tasks of trajectory forecasting and human pose prediction. For each of the two tasks, specialized models have been developed. Combining these models for holistic human motion prediction is non-trivial, and recent methods have struggled to compete on established benchmarks for individual tasks. To address this, we propose a simple yet effective transformer-based model for human motion prediction. The model employs a stack of self-attention modules to effectively capture both spatial dependencies within a pose and temporal relationships across a motion sequence. This simple, streamlined, end-to-end model is sufficiently versatile to handle pose-only, trajectory-only, and combined prediction tasks without task-specific modifications. We demonstrate that this approach achieves state-of-the-art results across all tasks through extensive experiments on a wide range of benchmark datasets, including Human3.6M, AMASS, ETH-UCY, and 3DPW.
title SimpliHuMoN: Simplifying Human Motion Prediction
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2603.04399