GP-VLS: A general-purpose vision language model for surgery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schmidgall, Samuel, Cho, Joseph, Zakka, Cyril, Hiesinger, William
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914903386750976
author Schmidgall, Samuel
Cho, Joseph
Zakka, Cyril
Hiesinger, William
author_facet Schmidgall, Samuel
Cho, Joseph
Zakka, Cyril
Hiesinger, William
contents Surgery requires comprehensive medical knowledge, visual assessment skills, and procedural expertise. While recent surgical AI models have focused on solving task-specific problems, there is a need for general-purpose systems that can understand surgical scenes and interact through natural language. This paper introduces GP-VLS, a general-purpose vision language model for surgery that integrates medical and surgical knowledge with visual scene understanding. For comprehensively evaluating general-purpose surgical models, we propose SurgiQual, which evaluates across medical and surgical knowledge benchmarks as well as surgical vision-language questions. To train GP-VLS, we develop six new datasets spanning medical knowledge, surgical textbooks, and vision-language pairs for tasks like phase recognition and tool identification. We show that GP-VLS significantly outperforms existing open- and closed-source models on surgical vision-language tasks, with 8-21% improvements in accuracy across SurgiQual benchmarks. GP-VLS also demonstrates strong performance on medical and surgical knowledge tests compared to open-source alternatives. Overall, GP-VLS provides an open-source foundation for developing AI assistants to support surgeons across a wide range of tasks and scenarios. The code and data for this work is publicly available at gpvls-surgery-vlm.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2407_19305
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GP-VLS: A general-purpose vision language model for surgery
Schmidgall, Samuel
Cho, Joseph
Zakka, Cyril
Hiesinger, William
Computer Vision and Pattern Recognition
Machine Learning
Tissues and Organs
Surgery requires comprehensive medical knowledge, visual assessment skills, and procedural expertise. While recent surgical AI models have focused on solving task-specific problems, there is a need for general-purpose systems that can understand surgical scenes and interact through natural language. This paper introduces GP-VLS, a general-purpose vision language model for surgery that integrates medical and surgical knowledge with visual scene understanding. For comprehensively evaluating general-purpose surgical models, we propose SurgiQual, which evaluates across medical and surgical knowledge benchmarks as well as surgical vision-language questions. To train GP-VLS, we develop six new datasets spanning medical knowledge, surgical textbooks, and vision-language pairs for tasks like phase recognition and tool identification. We show that GP-VLS significantly outperforms existing open- and closed-source models on surgical vision-language tasks, with 8-21% improvements in accuracy across SurgiQual benchmarks. GP-VLS also demonstrates strong performance on medical and surgical knowledge tests compared to open-source alternatives. Overall, GP-VLS provides an open-source foundation for developing AI assistants to support surgeons across a wide range of tasks and scenarios. The code and data for this work is publicly available at gpvls-surgery-vlm.github.io.
title GP-VLS: A general-purpose vision language model for surgery
topic Computer Vision and Pattern Recognition
Machine Learning
Tissues and Organs
url https://arxiv.org/abs/2407.19305