PhySU-Net: Long Temporal Context Transformer for rPPG with Self-Supervised Pre-training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Savic, Marko, Zhao, Guoying
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911779483811840
author Savic, Marko
Zhao, Guoying
author_facet Savic, Marko
Zhao, Guoying
contents Remote photoplethysmography (rPPG) is a promising technology that consists of contactless measuring of cardiac activity from facial videos. Most recent approaches utilize convolutional networks with limited temporal modeling capability or ignore long temporal context. Supervised rPPG methods are also severely limited by scarce data availability. In this work, we propose PhySU-Net, the first long spatial-temporal map rPPG transformer network and a self-supervised pre-training strategy that exploits unlabeled data to improve our model. Our strategy leverages traditional methods and image masking to provide pseudo-labels for self-supervised pre-training. Our model is tested on two public datasets (OBF and VIPL-HR) and shows superior performance in supervised training. Furthermore, we demonstrate that our self-supervised pre-training strategy further improves our model's performance by leveraging representations learned from unlabeled data.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11913
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PhySU-Net: Long Temporal Context Transformer for rPPG with Self-Supervised Pre-training
Savic, Marko
Zhao, Guoying
Computer Vision and Pattern Recognition
Remote photoplethysmography (rPPG) is a promising technology that consists of contactless measuring of cardiac activity from facial videos. Most recent approaches utilize convolutional networks with limited temporal modeling capability or ignore long temporal context. Supervised rPPG methods are also severely limited by scarce data availability. In this work, we propose PhySU-Net, the first long spatial-temporal map rPPG transformer network and a self-supervised pre-training strategy that exploits unlabeled data to improve our model. Our strategy leverages traditional methods and image masking to provide pseudo-labels for self-supervised pre-training. Our model is tested on two public datasets (OBF and VIPL-HR) and shows superior performance in supervised training. Furthermore, we demonstrate that our self-supervised pre-training strategy further improves our model's performance by leveraging representations learned from unlabeled data.
title PhySU-Net: Long Temporal Context Transformer for rPPG with Self-Supervised Pre-training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.11913