Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912525795196928 |
|---|---|
| author | Kallakurik, Uttej Humes, Edward Jonna, Rithvik Lin, Xiaomin Mohsenin, Tinoosh |
| author_facet | Kallakurik, Uttej Humes, Edward Jonna, Rithvik Lin, Xiaomin Mohsenin, Tinoosh |
| contents | Large Language Models (LLMs) have significant impact on the healthcare scenarios but remain prohibitively large for deployment in real-time, resource-constrained environments such as edge devices. In this work, we introduce a novel medical assistant system, optimized through our general-purpose compression framework, which tailors Large Language Models (LLMs) for deployment in specialized domains. By measuring neuron saliency on domain-specific data, our method can aggressively prune irrelevant neurons, reducing model size while preserving performance. Following pruning, we apply post-training quantization to further reduce the memory footprint, and evaluate the compressed model across medical benchmarks including MedMCQA, MedQA, and PubMedQA. We also deploy the 50\% compressed Gemma and the 67\% compressed LLaMA3 models on Jetson Orin Nano (18.7W peak) and Raspberry Pi 5 (6.3W peak), achieving real-time, energy-efficient inference under hardware constraints. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_11105 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation Kallakurik, Uttej Humes, Edward Jonna, Rithvik Lin, Xiaomin Mohsenin, Tinoosh Computation and Language Artificial Intelligence Hardware Architecture Systems and Control Large Language Models (LLMs) have significant impact on the healthcare scenarios but remain prohibitively large for deployment in real-time, resource-constrained environments such as edge devices. In this work, we introduce a novel medical assistant system, optimized through our general-purpose compression framework, which tailors Large Language Models (LLMs) for deployment in specialized domains. By measuring neuron saliency on domain-specific data, our method can aggressively prune irrelevant neurons, reducing model size while preserving performance. Following pruning, we apply post-training quantization to further reduce the memory footprint, and evaluate the compressed model across medical benchmarks including MedMCQA, MedQA, and PubMedQA. We also deploy the 50\% compressed Gemma and the 67\% compressed LLaMA3 models on Jetson Orin Nano (18.7W peak) and Raspberry Pi 5 (6.3W peak), achieving real-time, energy-efficient inference under hardware constraints. |
| title | Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation |
| topic | Computation and Language Artificial Intelligence Hardware Architecture Systems and Control |
| url | https://arxiv.org/abs/2506.11105 |