C3T: Cross-modal Transfer Through Time for Sensor-based Human Activity Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kamboj, Abhi, Nguyen, Anh Duy, Do, Minh N.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909681986830336
author Kamboj, Abhi
Nguyen, Anh Duy
Do, Minh N.
author_facet Kamboj, Abhi
Nguyen, Anh Duy
Do, Minh N.
contents In order to unlock the potential of diverse sensors, we investigate a method to transfer knowledge between time-series modalities using a multimodal \textit{temporal} representation space for Human Activity Recognition (HAR). Specifically, we explore the setting where the modality used in testing has no labeled data during training, which we refer to as Unsupervised Modality Adaptation (UMA). We categorize existing UMA approaches as Student-Teacher or Contrastive Alignment methods. These methods typically compress continuous-time data samples into single latent vectors during alignment, inhibiting their ability to transfer temporal information through real-world temporal distortions. To address this, we introduce Cross-modal Transfer Through Time (C3T), which preserves temporal information during alignment to handle dynamic sensor data better. C3T achieves this by aligning a set of temporal latent vectors across sensing modalities. Our extensive experiments on various camera+IMU datasets demonstrate that C3T outperforms existing methods in UMA by at least 8% in accuracy and shows superior robustness to temporal distortions such as time-shift, misalignment, and dilation. Our findings suggest that C3T has significant potential for developing generalizable models for time-series sensor data, opening new avenues for various multimodal applications.
format Preprint
id arxiv_https___arxiv_org_abs_2407_16803
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle C3T: Cross-modal Transfer Through Time for Sensor-based Human Activity Recognition
Kamboj, Abhi
Nguyen, Anh Duy
Do, Minh N.
Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Signal Processing
In order to unlock the potential of diverse sensors, we investigate a method to transfer knowledge between time-series modalities using a multimodal \textit{temporal} representation space for Human Activity Recognition (HAR). Specifically, we explore the setting where the modality used in testing has no labeled data during training, which we refer to as Unsupervised Modality Adaptation (UMA). We categorize existing UMA approaches as Student-Teacher or Contrastive Alignment methods. These methods typically compress continuous-time data samples into single latent vectors during alignment, inhibiting their ability to transfer temporal information through real-world temporal distortions. To address this, we introduce Cross-modal Transfer Through Time (C3T), which preserves temporal information during alignment to handle dynamic sensor data better. C3T achieves this by aligning a set of temporal latent vectors across sensing modalities. Our extensive experiments on various camera+IMU datasets demonstrate that C3T outperforms existing methods in UMA by at least 8% in accuracy and shows superior robustness to temporal distortions such as time-shift, misalignment, and dilation. Our findings suggest that C3T has significant potential for developing generalizable models for time-series sensor data, opening new avenues for various multimodal applications.
title C3T: Cross-modal Transfer Through Time for Sensor-based Human Activity Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Signal Processing
url https://arxiv.org/abs/2407.16803