Combo-Gait: Unified Transformer Framework for Multi-Modal Gait Recognition and Attribute Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhao-Yang, Shao, Zhimin, Nanduri, Anirudh, Pal, Basudha, McDaniel, Laura, Chen, Jieneng, Chellappa, Rama
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910155298308096
author Wang, Zhao-Yang
Shao, Zhimin
Nanduri, Anirudh
Pal, Basudha
McDaniel, Laura
Chen, Jieneng
Chellappa, Rama
author_facet Wang, Zhao-Yang
Shao, Zhimin
Nanduri, Anirudh
Pal, Basudha
McDaniel, Laura
Chen, Jieneng
Chellappa, Rama
contents Gait recognition is an important biometric for human identification at a distance, particularly under low-resolution or unconstrained environments. Current works typically focus on either 2D representations (e.g., silhouettes and skeletons) or 3D representations (e.g., meshes and SMPLs), but relying on a single modality often fails to capture the full geometric and dynamic complexity of human walking patterns. In this paper, we propose a multi-modal and multi-task framework that combines 2D temporal silhouettes with 3D SMPL features for robust gait analysis. Beyond identification, we introduce a multitask learning strategy that jointly performs gait recognition and human attribute estimation, including age, body mass index (BMI), and gender. A unified transformer is employed to effectively fuse multi-modal gait features and better learn attribute-related representations, while preserving discriminative identity cues. Extensive experiments on the large-scale BRIAR datasets, collected under challenging conditions such as long-range distances (up to 1 km) and extreme pitch angles (up to 50°), demonstrate that our approach outperforms state-of-the-art methods in gait recognition and provides accurate human attribute estimation. These results highlight the promise of multi-modal and multitask learning for advancing gait-based human understanding in real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10417
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Combo-Gait: Unified Transformer Framework for Multi-Modal Gait Recognition and Attribute Analysis
Wang, Zhao-Yang
Shao, Zhimin
Nanduri, Anirudh
Pal, Basudha
McDaniel, Laura
Chen, Jieneng
Chellappa, Rama
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Gait recognition is an important biometric for human identification at a distance, particularly under low-resolution or unconstrained environments. Current works typically focus on either 2D representations (e.g., silhouettes and skeletons) or 3D representations (e.g., meshes and SMPLs), but relying on a single modality often fails to capture the full geometric and dynamic complexity of human walking patterns. In this paper, we propose a multi-modal and multi-task framework that combines 2D temporal silhouettes with 3D SMPL features for robust gait analysis. Beyond identification, we introduce a multitask learning strategy that jointly performs gait recognition and human attribute estimation, including age, body mass index (BMI), and gender. A unified transformer is employed to effectively fuse multi-modal gait features and better learn attribute-related representations, while preserving discriminative identity cues. Extensive experiments on the large-scale BRIAR datasets, collected under challenging conditions such as long-range distances (up to 1 km) and extreme pitch angles (up to 50°), demonstrate that our approach outperforms state-of-the-art methods in gait recognition and provides accurate human attribute estimation. These results highlight the promise of multi-modal and multitask learning for advancing gait-based human understanding in real-world scenarios.
title Combo-Gait: Unified Transformer Framework for Multi-Modal Gait Recognition and Attribute Analysis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.10417