Specialized Foundation Models for Intelligent Operating Rooms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Özsoy, Ege, Pellegrini, Chantal, Bani-Harouni, David, Yuan, Kun, Keicher, Matthias, Navab, Nassir
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908433181048832
author Özsoy, Ege
Pellegrini, Chantal
Bani-Harouni, David
Yuan, Kun
Keicher, Matthias
Navab, Nassir
author_facet Özsoy, Ege
Pellegrini, Chantal
Bani-Harouni, David
Yuan, Kun
Keicher, Matthias
Navab, Nassir
contents Surgical procedures unfold in complex environments demanding coordination between surgical teams, tools, imaging and increasingly, intelligent robotic systems. Ensuring safety and efficiency in ORs of the future requires intelligent systems, like surgical robots, smart instruments and digital copilots, capable of understanding complex activities and hazards of surgeries. Yet, existing computational approaches, lack the breadth, and generalization needed for comprehensive OR understanding. We introduce ORQA, a multimodal foundation model unifying visual, auditory, and structured data for holistic surgical understanding. ORQA's question-answering framework empowers diverse tasks, serving as an intelligence core for a broad spectrum of surgical technologies. We benchmark ORQA against generalist vision-language models, including ChatGPT and Gemini, and show that while they struggle to perceive surgical scenes, ORQA delivers substantially stronger, consistent performance. Recognizing the extensive range of deployment settings across clinical practice, we design, and release a family of smaller ORQA models tailored to different computational requirements. This work establishes a foundation for the next wave of intelligent surgical solutions, enabling surgical teams and medical technology providers to create smarter and safer operating rooms.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12890
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Specialized Foundation Models for Intelligent Operating Rooms
Özsoy, Ege
Pellegrini, Chantal
Bani-Harouni, David
Yuan, Kun
Keicher, Matthias
Navab, Nassir
Computer Vision and Pattern Recognition
Surgical procedures unfold in complex environments demanding coordination between surgical teams, tools, imaging and increasingly, intelligent robotic systems. Ensuring safety and efficiency in ORs of the future requires intelligent systems, like surgical robots, smart instruments and digital copilots, capable of understanding complex activities and hazards of surgeries. Yet, existing computational approaches, lack the breadth, and generalization needed for comprehensive OR understanding. We introduce ORQA, a multimodal foundation model unifying visual, auditory, and structured data for holistic surgical understanding. ORQA's question-answering framework empowers diverse tasks, serving as an intelligence core for a broad spectrum of surgical technologies. We benchmark ORQA against generalist vision-language models, including ChatGPT and Gemini, and show that while they struggle to perceive surgical scenes, ORQA delivers substantially stronger, consistent performance. Recognizing the extensive range of deployment settings across clinical practice, we design, and release a family of smaller ORQA models tailored to different computational requirements. This work establishes a foundation for the next wave of intelligent surgical solutions, enabling surgical teams and medical technology providers to create smarter and safer operating rooms.
title Specialized Foundation Models for Intelligent Operating Rooms
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.12890