Explicit Time-Frequency Dynamics for Skeleton-Based Gait Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ko, Seoyeon, Song, Yeojin, Chung, Egene, Quagliato, Luca, Lee, Taeyong, Noh, Junhyug
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911564978716672
author Ko, Seoyeon
Song, Yeojin
Chung, Egene
Quagliato, Luca
Lee, Taeyong
Noh, Junhyug
author_facet Ko, Seoyeon
Song, Yeojin
Chung, Egene
Quagliato, Luca
Lee, Taeyong
Noh, Junhyug
contents Skeleton-based gait recognizers excel at modeling spatial configurations but often underuse explicit motion dynamics that are crucial under appearance changes. We introduce a plug-and-play Wavelet Feature Stream that augments any skeleton backbone with time-frequency dynamics of joint velocities. Concretely, per-joint velocity sequences are transformed by the continuous wavelet transform (CWT) into multi-scale scalograms, from which a lightweight multi-scale CNN learns discriminative dynamic cues. The resulting descriptor is fused with the backbone representation for classification, requiring no changes to the backbone architecture or additional supervision. Across CASIA-B, the proposed stream delivers consistent gains on strong skeleton backbones (e.g., GaitMixer, GaitFormer, GaitGraph) and establishes a new skeleton-based state of the art when attached to GaitMixer. The improvements are especially pronounced under covariate shifts such as carrying bags (BG) and wearing coats (CL), highlighting the complementarity of explicit time-frequency modeling and standard spatio-temporal encoders.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03002
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Explicit Time-Frequency Dynamics for Skeleton-Based Gait Recognition
Ko, Seoyeon
Song, Yeojin
Chung, Egene
Quagliato, Luca
Lee, Taeyong
Noh, Junhyug
Computer Vision and Pattern Recognition
Skeleton-based gait recognizers excel at modeling spatial configurations but often underuse explicit motion dynamics that are crucial under appearance changes. We introduce a plug-and-play Wavelet Feature Stream that augments any skeleton backbone with time-frequency dynamics of joint velocities. Concretely, per-joint velocity sequences are transformed by the continuous wavelet transform (CWT) into multi-scale scalograms, from which a lightweight multi-scale CNN learns discriminative dynamic cues. The resulting descriptor is fused with the backbone representation for classification, requiring no changes to the backbone architecture or additional supervision. Across CASIA-B, the proposed stream delivers consistent gains on strong skeleton backbones (e.g., GaitMixer, GaitFormer, GaitGraph) and establishes a new skeleton-based state of the art when attached to GaitMixer. The improvements are especially pronounced under covariate shifts such as carrying bags (BG) and wearing coats (CL), highlighting the complementarity of explicit time-frequency modeling and standard spatio-temporal encoders.
title Explicit Time-Frequency Dynamics for Skeleton-Based Gait Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.03002