On-device Large Multi-modal Agent for Human Activity Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Siam, Md Shakhrul Iman, Showmik, Ishtiaque Ahmed, Song, Guanqun, Zhu, Ting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909973826502656
author Siam, Md Shakhrul Iman
Showmik, Ishtiaque Ahmed
Song, Guanqun
Zhu, Ting
author_facet Siam, Md Shakhrul Iman
Showmik, Ishtiaque Ahmed
Song, Guanqun
Zhu, Ting
contents Human Activity Recognition (HAR) has been an active area of research, with applications ranging from healthcare to smart environments. The recent advancements in Large Language Models (LLMs) have opened new possibilities to leverage their capabilities in HAR, enabling not just activity classification but also interpretability and human-like interaction. In this paper, we present a Large Multi-Modal Agent designed for HAR, which integrates the power of LLMs to enhance both performance and user engagement. The proposed framework not only delivers activity classification but also bridges the gap between technical outputs and user-friendly insights through its reasoning and question-answering capabilities. We conduct extensive evaluations using widely adopted HAR datasets, including HHAR, Shoaib, Motionsense to assess the performance of our framework. The results demonstrate that our model achieves high classification accuracy comparable to state-of-the-art methods while significantly improving interpretability through its reasoning and Q&A capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19742
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On-device Large Multi-modal Agent for Human Activity Recognition
Siam, Md Shakhrul Iman
Showmik, Ishtiaque Ahmed
Song, Guanqun
Zhu, Ting
Machine Learning
Human Activity Recognition (HAR) has been an active area of research, with applications ranging from healthcare to smart environments. The recent advancements in Large Language Models (LLMs) have opened new possibilities to leverage their capabilities in HAR, enabling not just activity classification but also interpretability and human-like interaction. In this paper, we present a Large Multi-Modal Agent designed for HAR, which integrates the power of LLMs to enhance both performance and user engagement. The proposed framework not only delivers activity classification but also bridges the gap between technical outputs and user-friendly insights through its reasoning and question-answering capabilities. We conduct extensive evaluations using widely adopted HAR datasets, including HHAR, Shoaib, Motionsense to assess the performance of our framework. The results demonstrate that our model achieves high classification accuracy comparable to state-of-the-art methods while significantly improving interpretability through its reasoning and Q&A capabilities.
title On-device Large Multi-modal Agent for Human Activity Recognition
topic Machine Learning
url https://arxiv.org/abs/2512.19742