SETransformer: A Hybrid Attention-Based Architecture for Robust Human Activity Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Yunbo, Qin, Xukui, Gao, Yifan, Li, Xiang, Feng, Chengwei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915304299298816
author Liu, Yunbo
Qin, Xukui
Gao, Yifan
Li, Xiang
Feng, Chengwei
author_facet Liu, Yunbo
Qin, Xukui
Gao, Yifan
Li, Xiang
Feng, Chengwei
contents Human Activity Recognition (HAR) using wearable sensor data has become a central task in mobile computing, healthcare, and human-computer interaction. Despite the success of traditional deep learning models such as CNNs and RNNs, they often struggle to capture long-range temporal dependencies and contextual relevance across multiple sensor channels. To address these limitations, we propose SETransformer, a hybrid deep neural architecture that combines Transformer-based temporal modeling with channel-wise squeeze-and-excitation (SE) attention and a learnable temporal attention pooling mechanism. The model takes raw triaxial accelerometer data as input and leverages global self-attention to capture activity-specific motion dynamics over extended time windows, while adaptively emphasizing informative sensor channels and critical time steps. We evaluate SETransformer on the WISDM dataset and demonstrate that it significantly outperforms conventional models including LSTM, GRU, BiLSTM, and CNN baselines. The proposed model achieves a validation accuracy of 84.68\% and a macro F1-score of 84.64\%, surpassing all baseline architectures by a notable margin. Our results show that SETransformer is a competitive and interpretable solution for real-world HAR tasks, with strong potential for deployment in mobile and ubiquitous sensing applications.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19369
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SETransformer: A Hybrid Attention-Based Architecture for Robust Human Activity Recognition
Liu, Yunbo
Qin, Xukui
Gao, Yifan
Li, Xiang
Feng, Chengwei
Machine Learning
Artificial Intelligence
Human Activity Recognition (HAR) using wearable sensor data has become a central task in mobile computing, healthcare, and human-computer interaction. Despite the success of traditional deep learning models such as CNNs and RNNs, they often struggle to capture long-range temporal dependencies and contextual relevance across multiple sensor channels. To address these limitations, we propose SETransformer, a hybrid deep neural architecture that combines Transformer-based temporal modeling with channel-wise squeeze-and-excitation (SE) attention and a learnable temporal attention pooling mechanism. The model takes raw triaxial accelerometer data as input and leverages global self-attention to capture activity-specific motion dynamics over extended time windows, while adaptively emphasizing informative sensor channels and critical time steps. We evaluate SETransformer on the WISDM dataset and demonstrate that it significantly outperforms conventional models including LSTM, GRU, BiLSTM, and CNN baselines. The proposed model achieves a validation accuracy of 84.68\% and a macro F1-score of 84.64\%, surpassing all baseline architectures by a notable margin. Our results show that SETransformer is a competitive and interpretable solution for real-world HAR tasks, with strong potential for deployment in mobile and ubiquitous sensing applications.
title SETransformer: A Hybrid Attention-Based Architecture for Robust Human Activity Recognition
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.19369