(LiFT) Lightweight Fitness Transformer: A language-vision model for Remote Monitoring of Physical Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Postlmayr, A., Cosman, P., Dey, S.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915331892576256
author Postlmayr, A.
Cosman, P.
Dey, S.
author_facet Postlmayr, A.
Cosman, P.
Dey, S.
contents We introduce a fitness tracking system that enables remote monitoring for exercises using only a RGB smartphone camera, making fitness tracking more private, scalable, and cost effective. Although prior work explored automated exercise supervision, existing models are either too limited in exercise variety or too complex for real-world deployment. Prior approaches typically focus on a small set of exercises and fail to generalize across diverse movements. In contrast, we develop a robust, multitask motion analysis model capable of performing exercise detection and repetition counting across hundreds of exercises, a scale far beyond previous methods. We overcome previous data limitations by assembling a large-scale fitness dataset, Olympia covering more than 1,900 exercises. To our knowledge, our vision-language model is the first that can perform multiple tasks on skeletal fitness data. On Olympia, our model can detect exercises with 76.5% accuracy and count repetitions with 85.3% off-by-one accuracy, using only RGB video. By presenting a single vision-language transformer model for both exercise identification and rep counting, we take a significant step toward democratizing AI-powered fitness tracking.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06480
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle (LiFT) Lightweight Fitness Transformer: A language-vision model for Remote Monitoring of Physical Training
Postlmayr, A.
Cosman, P.
Dey, S.
Computer Vision and Pattern Recognition
We introduce a fitness tracking system that enables remote monitoring for exercises using only a RGB smartphone camera, making fitness tracking more private, scalable, and cost effective. Although prior work explored automated exercise supervision, existing models are either too limited in exercise variety or too complex for real-world deployment. Prior approaches typically focus on a small set of exercises and fail to generalize across diverse movements. In contrast, we develop a robust, multitask motion analysis model capable of performing exercise detection and repetition counting across hundreds of exercises, a scale far beyond previous methods. We overcome previous data limitations by assembling a large-scale fitness dataset, Olympia covering more than 1,900 exercises. To our knowledge, our vision-language model is the first that can perform multiple tasks on skeletal fitness data. On Olympia, our model can detect exercises with 76.5% accuracy and count repetitions with 85.3% off-by-one accuracy, using only RGB video. By presenting a single vision-language transformer model for both exercise identification and rep counting, we take a significant step toward democratizing AI-powered fitness tracking.
title (LiFT) Lightweight Fitness Transformer: A language-vision model for Remote Monitoring of Physical Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.06480