Saved in:
Bibliographic Details
Main Authors: Ko, Seoyeon, Song, Yeojin, Chung, Egene, Quagliato, Luca, Lee, Taeyong, Noh, Junhyug
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.03002
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911564978716672
author Ko, Seoyeon
Song, Yeojin
Chung, Egene
Quagliato, Luca
Lee, Taeyong
Noh, Junhyug
author_facet Ko, Seoyeon
Song, Yeojin
Chung, Egene
Quagliato, Luca
Lee, Taeyong
Noh, Junhyug
contents Skeleton-based gait recognizers excel at modeling spatial configurations but often underuse explicit motion dynamics that are crucial under appearance changes. We introduce a plug-and-play Wavelet Feature Stream that augments any skeleton backbone with time-frequency dynamics of joint velocities. Concretely, per-joint velocity sequences are transformed by the continuous wavelet transform (CWT) into multi-scale scalograms, from which a lightweight multi-scale CNN learns discriminative dynamic cues. The resulting descriptor is fused with the backbone representation for classification, requiring no changes to the backbone architecture or additional supervision. Across CASIA-B, the proposed stream delivers consistent gains on strong skeleton backbones (e.g., GaitMixer, GaitFormer, GaitGraph) and establishes a new skeleton-based state of the art when attached to GaitMixer. The improvements are especially pronounced under covariate shifts such as carrying bags (BG) and wearing coats (CL), highlighting the complementarity of explicit time-frequency modeling and standard spatio-temporal encoders.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03002
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Explicit Time-Frequency Dynamics for Skeleton-Based Gait Recognition
Ko, Seoyeon
Song, Yeojin
Chung, Egene
Quagliato, Luca
Lee, Taeyong
Noh, Junhyug
Computer Vision and Pattern Recognition
Skeleton-based gait recognizers excel at modeling spatial configurations but often underuse explicit motion dynamics that are crucial under appearance changes. We introduce a plug-and-play Wavelet Feature Stream that augments any skeleton backbone with time-frequency dynamics of joint velocities. Concretely, per-joint velocity sequences are transformed by the continuous wavelet transform (CWT) into multi-scale scalograms, from which a lightweight multi-scale CNN learns discriminative dynamic cues. The resulting descriptor is fused with the backbone representation for classification, requiring no changes to the backbone architecture or additional supervision. Across CASIA-B, the proposed stream delivers consistent gains on strong skeleton backbones (e.g., GaitMixer, GaitFormer, GaitGraph) and establishes a new skeleton-based state of the art when attached to GaitMixer. The improvements are especially pronounced under covariate shifts such as carrying bags (BG) and wearing coats (CL), highlighting the complementarity of explicit time-frequency modeling and standard spatio-temporal encoders.
title Explicit Time-Frequency Dynamics for Skeleton-Based Gait Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.03002