CMD-HAR: Cross-Modal Disentanglement for Wearable Human Activity Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Ying, Li, Siyao, Jiang, Yixuan, Xiao, Hang, Long, Jingxi, Tang, Haotian, Liu, Hanyu, Li, Chao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914306988179456
author Yu, Ying
Li, Siyao
Jiang, Yixuan
Xiao, Hang
Long, Jingxi
Tang, Haotian
Liu, Hanyu
Li, Chao
author_facet Yu, Ying
Li, Siyao
Jiang, Yixuan
Xiao, Hang
Long, Jingxi
Tang, Haotian
Liu, Hanyu
Li, Chao
contents Human Activity Recognition (HAR) is a fundamental technology for numerous human - centered intelligent applications. Although deep learning methods have been utilized to accelerate feature extraction, issues such as multimodal data mixing, activity heterogeneity, and complex model deployment remain largely unresolved. The aim of this paper is to address issues such as multimodal data mixing, activity heterogeneity, and complex model deployment in sensor-based human activity recognition. We propose a spatiotemporal attention modal decomposition alignment fusion strategy to tackle the problem of the mixed distribution of sensor data. Key discriminative features of activities are captured through cross-modal spatio-temporal disentangled representation, and gradient modulation is combined to alleviate data heterogeneity. In addition, a wearable deployment simulation system is constructed. We conducted experiments on a large number of public datasets, demonstrating the effectiveness of the model.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21843
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CMD-HAR: Cross-Modal Disentanglement for Wearable Human Activity Recognition
Yu, Ying
Li, Siyao
Jiang, Yixuan
Xiao, Hang
Long, Jingxi
Tang, Haotian
Liu, Hanyu
Li, Chao
Computer Vision and Pattern Recognition
Artificial Intelligence
Human Activity Recognition (HAR) is a fundamental technology for numerous human - centered intelligent applications. Although deep learning methods have been utilized to accelerate feature extraction, issues such as multimodal data mixing, activity heterogeneity, and complex model deployment remain largely unresolved. The aim of this paper is to address issues such as multimodal data mixing, activity heterogeneity, and complex model deployment in sensor-based human activity recognition. We propose a spatiotemporal attention modal decomposition alignment fusion strategy to tackle the problem of the mixed distribution of sensor data. Key discriminative features of activities are captured through cross-modal spatio-temporal disentangled representation, and gradient modulation is combined to alleviate data heterogeneity. In addition, a wearable deployment simulation system is constructed. We conducted experiments on a large number of public datasets, demonstrating the effectiveness of the model.
title CMD-HAR: Cross-Modal Disentanglement for Wearable Human Activity Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.21843