Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Zongxian, Qian, Jiayu, Tan, Kay Chen, Wong, Hau-San, Chen, Yulong, Zhang, Haoyu, Huang, Zhi-An
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2504.12334
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913108488880128
author Yang, Zongxian
Qian, Jiayu
Tan, Kay Chen
Wong, Hau-San
Chen, Yulong
Zhang, Haoyu
Huang, Zhi-An
author_facet Yang, Zongxian
Qian, Jiayu
Tan, Kay Chen
Wong, Hau-San
Chen, Yulong
Zhang, Haoyu
Huang, Zhi-An
contents Large language models (LLMs) face significant challenges in specialized biomedical tasks due to the inherent complexity of medical reasoning and the sensitive nature of clinical data. Existing LLMs often struggle with intricate medical terminology and the need for accurate clinical insights, leading to performance reduction when quantized for resource-constrained deployment. To address these issues, we propose Quantized Medical Tree of Thought (QM-ToT), a path-based reasoning framework. QM-ToT leverages a Tree of Thought (ToT) reasoning approach to decompose complex medical problems into manageable subtasks, coupled with evaluator assessment layers. This framework facilitates substantial performance improvements in INT4-quantized models on the challenging MedQAUSMLE dataset. Specifically, we demonstrate a remarkable accuracy increase from 34% to 50% for the LLaMA2-70b model and from 58.77% to 69.49% for LLaMA-3.1-8b. Besides, we also proposed an effect data distillation method based on ToT. Compared to the traditional distillation method, we achieved an improvement of 86. 27% while using only 3.9% of the data.This work, for the first time, showcases the potential of ToT to significantly enhance performance on complex biomedical tasks, establishing a crucial foundation for future advances in deploying high-performing quantized LLM in resource-limited medical settings.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12334
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model
Yang, Zongxian
Qian, Jiayu
Tan, Kay Chen
Wong, Hau-San
Chen, Yulong
Zhang, Haoyu
Huang, Zhi-An
Computation and Language
Large language models (LLMs) face significant challenges in specialized biomedical tasks due to the inherent complexity of medical reasoning and the sensitive nature of clinical data. Existing LLMs often struggle with intricate medical terminology and the need for accurate clinical insights, leading to performance reduction when quantized for resource-constrained deployment. To address these issues, we propose Quantized Medical Tree of Thought (QM-ToT), a path-based reasoning framework. QM-ToT leverages a Tree of Thought (ToT) reasoning approach to decompose complex medical problems into manageable subtasks, coupled with evaluator assessment layers. This framework facilitates substantial performance improvements in INT4-quantized models on the challenging MedQAUSMLE dataset. Specifically, we demonstrate a remarkable accuracy increase from 34% to 50% for the LLaMA2-70b model and from 58.77% to 69.49% for LLaMA-3.1-8b. Besides, we also proposed an effect data distillation method based on ToT. Compared to the traditional distillation method, we achieved an improvement of 86. 27% while using only 3.9% of the data.This work, for the first time, showcases the potential of ToT to significantly enhance performance on complex biomedical tasks, establishing a crucial foundation for future advances in deploying high-performing quantized LLM in resource-limited medical settings.
title QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model
topic Computation and Language
url https://arxiv.org/abs/2504.12334