SensorLM: Learning the Language of Wearable Sensors
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918054638649344 |
|---|---|
| author | Zhang, Yuwei Ayush, Kumar Qiao, Siyuan Heydari, A. Ali Narayanswamy, Girish Xu, Maxwell A. Metwally, Ahmed A. Xu, Shawn Garrison, Jake Xu, Xuhai Althoff, Tim Liu, Yun Kohli, Pushmeet Zhan, Jiening Malhotra, Mark Patel, Shwetak Mascolo, Cecilia Liu, Xin McDuff, Daniel Yang, Yuzhe |
| author_facet | Zhang, Yuwei Ayush, Kumar Qiao, Siyuan Heydari, A. Ali Narayanswamy, Girish Xu, Maxwell A. Metwally, Ahmed A. Xu, Shawn Garrison, Jake Xu, Xuhai Althoff, Tim Liu, Yun Kohli, Pushmeet Zhan, Jiening Malhotra, Mark Patel, Shwetak Mascolo, Cecilia Liu, Xin McDuff, Daniel Yang, Yuzhe |
| contents | We present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and interpreting sensor data with language remains challenging due to the lack of paired, richly annotated sensor-text descriptions in uncurated, real-world wearable data. We introduce a hierarchical caption generation pipeline designed to capture statistical, structural, and semantic information from sensor data. This approach enabled the curation of the largest sensor-language dataset to date, comprising over 59.7 million hours of data from more than 103,000 people. Furthermore, SensorLM extends prominent multimodal pretraining architectures (e.g., CLIP, CoCa) and recovers them as specific variants within a generic architecture. Extensive experiments on real-world tasks in human activity analysis and healthcare verify the superior performance of SensorLM over state-of-the-art in zero-shot recognition, few-shot learning, and cross-modal retrieval. SensorLM also demonstrates intriguing capabilities including scaling behaviors, label efficiency, sensor captioning, and zero-shot generalization to unseen tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_09108 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SensorLM: Learning the Language of Wearable Sensors Zhang, Yuwei Ayush, Kumar Qiao, Siyuan Heydari, A. Ali Narayanswamy, Girish Xu, Maxwell A. Metwally, Ahmed A. Xu, Shawn Garrison, Jake Xu, Xuhai Althoff, Tim Liu, Yun Kohli, Pushmeet Zhan, Jiening Malhotra, Mark Patel, Shwetak Mascolo, Cecilia Liu, Xin McDuff, Daniel Yang, Yuzhe Machine Learning Artificial Intelligence Computation and Language We present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and interpreting sensor data with language remains challenging due to the lack of paired, richly annotated sensor-text descriptions in uncurated, real-world wearable data. We introduce a hierarchical caption generation pipeline designed to capture statistical, structural, and semantic information from sensor data. This approach enabled the curation of the largest sensor-language dataset to date, comprising over 59.7 million hours of data from more than 103,000 people. Furthermore, SensorLM extends prominent multimodal pretraining architectures (e.g., CLIP, CoCa) and recovers them as specific variants within a generic architecture. Extensive experiments on real-world tasks in human activity analysis and healthcare verify the superior performance of SensorLM over state-of-the-art in zero-shot recognition, few-shot learning, and cross-modal retrieval. SensorLM also demonstrates intriguing capabilities including scaling behaviors, label efficiency, sensor captioning, and zero-shot generalization to unseen tasks. |
| title | SensorLM: Learning the Language of Wearable Sensors |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2506.09108 |