Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haresamudram, Harish, Beedu, Apoorva, Rabbi, Mashfiqui, Saha, Sankalita, Essa, Irfan, Ploetz, Thomas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929469262921728
author Haresamudram, Harish
Beedu, Apoorva
Rabbi, Mashfiqui
Saha, Sankalita
Essa, Irfan
Ploetz, Thomas
author_facet Haresamudram, Harish
Beedu, Apoorva
Rabbi, Mashfiqui
Saha, Sankalita
Essa, Irfan
Ploetz, Thomas
contents Cross-modal contrastive pre-training between natural language and other modalities, e.g., vision and audio, has demonstrated astonishing performance and effectiveness across a diverse variety of tasks and domains. In this paper, we investigate whether such natural language supervision can be used for wearable sensor based Human Activity Recognition (HAR), and discover that-surprisingly-it performs substantially worse than standard end-to-end training and self-supervision. We identify the primary causes for this as: sensor heterogeneity and the lack of rich, diverse text descriptions of activities. To mitigate their impact, we also develop strategies and assess their effectiveness through an extensive experimental evaluation. These strategies lead to significant increases in activity recognition, bringing performance closer to supervised and self-supervised training, while also enabling the recognition of unseen activities and cross modal retrieval of videos. Overall, our work paves the way for better sensor-language learning, ultimately leading to the development of foundational models for HAR using wearables.
format Preprint
id arxiv_https___arxiv_org_abs_2408_12023
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
Haresamudram, Harish
Beedu, Apoorva
Rabbi, Mashfiqui
Saha, Sankalita
Essa, Irfan
Ploetz, Thomas
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Cross-modal contrastive pre-training between natural language and other modalities, e.g., vision and audio, has demonstrated astonishing performance and effectiveness across a diverse variety of tasks and domains. In this paper, we investigate whether such natural language supervision can be used for wearable sensor based Human Activity Recognition (HAR), and discover that-surprisingly-it performs substantially worse than standard end-to-end training and self-supervision. We identify the primary causes for this as: sensor heterogeneity and the lack of rich, diverse text descriptions of activities. To mitigate their impact, we also develop strategies and assess their effectiveness through an extensive experimental evaluation. These strategies lead to significant increases in activity recognition, bringing performance closer to supervised and self-supervised training, while also enabling the recognition of unseen activities and cross modal retrieval of videos. Overall, our work paves the way for better sensor-language learning, ultimately leading to the development of foundational models for HAR using wearables.
title Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.12023