PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xie, Yiping, Zhao, Bo, Dai, Mingtong, Zhou, Jian-Ping, Sun, Yue, Tan, Tao, Xie, Weicheng, Shen, Linlin, Yu, Zitong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911485187325952
author Xie, Yiping
Zhao, Bo
Dai, Mingtong
Zhou, Jian-Ping
Sun, Yue
Tan, Tao
Xie, Weicheng
Shen, Linlin
Yu, Zitong
author_facet Xie, Yiping
Zhao, Bo
Dai, Mingtong
Zhou, Jian-Ping
Sun, Yue
Tan, Tao
Xie, Weicheng
Shen, Linlin
Yu, Zitong
contents Remote photoplethysmography (rPPG) enables non-contact physiological measurement but remains highly susceptible to illumination changes, motion artifacts, and limited temporal modeling. Large Language Models (LLMs) excel at capturing long-range dependencies, offering a potential solution but struggle with the continuous, noise-sensitive nature of rPPG signals due to their text-centric design. To bridge this gap, we introduce the PhysLLM, a collaborative optimization framework that synergizes LLMs with domain-specific rPPG components. Specifically, the Text Prototype Guidance (TPG) strategy is proposed to establish cross-modal alignment by projecting hemodynamic features into LLM-interpretable semantic space, effectively bridging the representational gap between physiological signals and linguistic tokens. Besides, a novel Dual-Domain Stationary (DDS) Algorithm is proposed for resolving signal instability through adaptive time-frequency domain feature re-weighting. Finally, rPPG task-specific cues systematically inject physiological priors through physiological statistics, environmental contextual answering, and task description, leveraging cross-modal learning to integrate both visual and textual information, enabling dynamic adaptation to challenging scenarios like variable illumination and subject movements. Evaluation on four benchmark datasets, PhysLLM achieves state-of-the-art accuracy and robustness, demonstrating superior generalization across lighting variations and motion scenarios. The source code is available at https://github.com/Alex036225/PhysLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2505_03621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing
Xie, Yiping
Zhao, Bo
Dai, Mingtong
Zhou, Jian-Ping
Sun, Yue
Tan, Tao
Xie, Weicheng
Shen, Linlin
Yu, Zitong
Computer Vision and Pattern Recognition
Remote photoplethysmography (rPPG) enables non-contact physiological measurement but remains highly susceptible to illumination changes, motion artifacts, and limited temporal modeling. Large Language Models (LLMs) excel at capturing long-range dependencies, offering a potential solution but struggle with the continuous, noise-sensitive nature of rPPG signals due to their text-centric design. To bridge this gap, we introduce the PhysLLM, a collaborative optimization framework that synergizes LLMs with domain-specific rPPG components. Specifically, the Text Prototype Guidance (TPG) strategy is proposed to establish cross-modal alignment by projecting hemodynamic features into LLM-interpretable semantic space, effectively bridging the representational gap between physiological signals and linguistic tokens. Besides, a novel Dual-Domain Stationary (DDS) Algorithm is proposed for resolving signal instability through adaptive time-frequency domain feature re-weighting. Finally, rPPG task-specific cues systematically inject physiological priors through physiological statistics, environmental contextual answering, and task description, leveraging cross-modal learning to integrate both visual and textual information, enabling dynamic adaptation to challenging scenarios like variable illumination and subject movements. Evaluation on four benchmark datasets, PhysLLM achieves state-of-the-art accuracy and robustness, demonstrating superior generalization across lighting variations and motion scenarios. The source code is available at https://github.com/Alex036225/PhysLLM.
title PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.03621