Saved in:
Bibliographic Details
Main Authors: Jin, Liuyi, Gunawardena, Pasan, Haroon, Amran, Wang, Runzhi, Lee, Sangwoo, Stoleru, Radu, Middleton, Michael, Huo, Zepeng, Kim, Jeeeun, Moats, Jason
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.13078
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915622332399616
author Jin, Liuyi
Gunawardena, Pasan
Haroon, Amran
Wang, Runzhi
Lee, Sangwoo
Stoleru, Radu
Middleton, Michael
Huo, Zepeng
Kim, Jeeeun
Moats, Jason
author_facet Jin, Liuyi
Gunawardena, Pasan
Haroon, Amran
Wang, Runzhi
Lee, Sangwoo
Stoleru, Radu
Middleton, Michael
Huo, Zepeng
Kim, Jeeeun
Moats, Jason
contents Emergency Medical Technicians (EMTs) operate in high-pressure environments, making rapid, life-critical decisions under heavy cognitive and operational loads. We present EMSGlass, a smart-glasses system powered by EMSNet, the first multimodal multitask model for Emergency Medical Services (EMS), and EMSServe, a low-latency multimodal serving framework tailored to EMS scenarios. EMSNet integrates text, vital signs, and scene images to construct a unified real-time understanding of EMS incidents. Trained on real-world multimodal EMS datasets, EMSNet simultaneously supports up to five critical EMS tasks with superior accuracy compared to state-of-the-art unimodal baselines. Built on top of PyTorch, EMSServe introduces a modality-aware model splitter and a feature caching mechanism, achieving adaptive and efficient inference across heterogeneous hardware while addressing the challenge of asynchronous modality arrival in the field. By optimizing multimodal inference execution in EMS scenarios, EMSServe achieves 1.9x -- 11.7x speedup over direct PyTorch multimodal inference. A user study evaluation with six professional EMTs demonstrates that EMSGlass enhances real-time situational awareness, decision-making speed, and operational efficiency through intuitive on-glass interaction. In addition, qualitative insights from the user study provide actionable directions for extending EMSGlass toward next-generation AI-enabled EMS systems, bridging multimodal intelligence with real-world emergency response workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13078
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Smart-Glasses for Emergency Medical Services via Multimodal Multitask Learning
Jin, Liuyi
Gunawardena, Pasan
Haroon, Amran
Wang, Runzhi
Lee, Sangwoo
Stoleru, Radu
Middleton, Michael
Huo, Zepeng
Kim, Jeeeun
Moats, Jason
Machine Learning
Audio and Speech Processing
Image and Video Processing
Emergency Medical Technicians (EMTs) operate in high-pressure environments, making rapid, life-critical decisions under heavy cognitive and operational loads. We present EMSGlass, a smart-glasses system powered by EMSNet, the first multimodal multitask model for Emergency Medical Services (EMS), and EMSServe, a low-latency multimodal serving framework tailored to EMS scenarios. EMSNet integrates text, vital signs, and scene images to construct a unified real-time understanding of EMS incidents. Trained on real-world multimodal EMS datasets, EMSNet simultaneously supports up to five critical EMS tasks with superior accuracy compared to state-of-the-art unimodal baselines. Built on top of PyTorch, EMSServe introduces a modality-aware model splitter and a feature caching mechanism, achieving adaptive and efficient inference across heterogeneous hardware while addressing the challenge of asynchronous modality arrival in the field. By optimizing multimodal inference execution in EMS scenarios, EMSServe achieves 1.9x -- 11.7x speedup over direct PyTorch multimodal inference. A user study evaluation with six professional EMTs demonstrates that EMSGlass enhances real-time situational awareness, decision-making speed, and operational efficiency through intuitive on-glass interaction. In addition, qualitative insights from the user study provide actionable directions for extending EMSGlass toward next-generation AI-enabled EMS systems, bridging multimodal intelligence with real-world emergency response workflows.
title A Smart-Glasses for Emergency Medical Services via Multimodal Multitask Learning
topic Machine Learning
Audio and Speech Processing
Image and Video Processing
url https://arxiv.org/abs/2511.13078