Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kallakurik, Uttej, Humes, Edward, Jonna, Rithvik, Lin, Xiaomin, Mohsenin, Tinoosh
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912525795196928
author Kallakurik, Uttej
Humes, Edward
Jonna, Rithvik
Lin, Xiaomin
Mohsenin, Tinoosh
author_facet Kallakurik, Uttej
Humes, Edward
Jonna, Rithvik
Lin, Xiaomin
Mohsenin, Tinoosh
contents Large Language Models (LLMs) have significant impact on the healthcare scenarios but remain prohibitively large for deployment in real-time, resource-constrained environments such as edge devices. In this work, we introduce a novel medical assistant system, optimized through our general-purpose compression framework, which tailors Large Language Models (LLMs) for deployment in specialized domains. By measuring neuron saliency on domain-specific data, our method can aggressively prune irrelevant neurons, reducing model size while preserving performance. Following pruning, we apply post-training quantization to further reduce the memory footprint, and evaluate the compressed model across medical benchmarks including MedMCQA, MedQA, and PubMedQA. We also deploy the 50\% compressed Gemma and the 67\% compressed LLaMA3 models on Jetson Orin Nano (18.7W peak) and Raspberry Pi 5 (6.3W peak), achieving real-time, energy-efficient inference under hardware constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11105
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation
Kallakurik, Uttej
Humes, Edward
Jonna, Rithvik
Lin, Xiaomin
Mohsenin, Tinoosh
Computation and Language
Artificial Intelligence
Hardware Architecture
Systems and Control
Large Language Models (LLMs) have significant impact on the healthcare scenarios but remain prohibitively large for deployment in real-time, resource-constrained environments such as edge devices. In this work, we introduce a novel medical assistant system, optimized through our general-purpose compression framework, which tailors Large Language Models (LLMs) for deployment in specialized domains. By measuring neuron saliency on domain-specific data, our method can aggressively prune irrelevant neurons, reducing model size while preserving performance. Following pruning, we apply post-training quantization to further reduce the memory footprint, and evaluate the compressed model across medical benchmarks including MedMCQA, MedQA, and PubMedQA. We also deploy the 50\% compressed Gemma and the 67\% compressed LLaMA3 models on Jetson Orin Nano (18.7W peak) and Raspberry Pi 5 (6.3W peak), achieving real-time, energy-efficient inference under hardware constraints.
title Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation
topic Computation and Language
Artificial Intelligence
Hardware Architecture
Systems and Control
url https://arxiv.org/abs/2506.11105