Saved in:
Bibliographic Details
Main Authors: Choi, Tae-Min, Jeong, Tae Kyeong, Kim, Garam, Lee, Jaemin, Koh, Yeongyoon, Choi, In Cheul, Chung, Jae-Ho, Park, Jong Woong, Park, Juyoun
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.21339
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915639127441408
author Choi, Tae-Min
Jeong, Tae Kyeong
Kim, Garam
Lee, Jaemin
Koh, Yeongyoon
Choi, In Cheul
Chung, Jae-Ho
Park, Jong Woong
Park, Juyoun
author_facet Choi, Tae-Min
Jeong, Tae Kyeong
Kim, Garam
Lee, Jaemin
Koh, Yeongyoon
Choi, In Cheul
Chung, Jae-Ho
Park, Jong Woong
Park, Juyoun
contents Recent advances in multimodal large language models (LLMs) have highlighted their potential for medical and surgical applications. However, existing surgical datasets predominantly adopt a Visual Question Answering (VQA) format with heterogeneous taxonomies and lack support for pixel-level segmentation, limiting consistent evaluation and applicability. We present SurgMLLMBench, a unified multimodal benchmark explicitly designed for developing and evaluating interactive multimodal LLMs for surgical scene understanding, including the newly collected Micro-surgical Artificial Vascular anastomosIS (MAVIS) dataset. It integrates pixel-level instrument segmentation masks and structured VQA annotations across laparoscopic, robot-assisted, and micro-surgical domains under a unified taxonomy, enabling comprehensive evaluation beyond traditional VQA tasks and richer visual-conversational interactions. Extensive baseline experiments show that a single model trained on SurgMLLMBench achieves consistent performance across domains and generalizes effectively to unseen datasets. SurgMLLMBench will be publicly released as a robust resource to advance multimodal surgical AI research, supporting reproducible evaluation and development of interactive surgical reasoning models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21339
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
Choi, Tae-Min
Jeong, Tae Kyeong
Kim, Garam
Lee, Jaemin
Koh, Yeongyoon
Choi, In Cheul
Chung, Jae-Ho
Park, Jong Woong
Park, Juyoun
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in multimodal large language models (LLMs) have highlighted their potential for medical and surgical applications. However, existing surgical datasets predominantly adopt a Visual Question Answering (VQA) format with heterogeneous taxonomies and lack support for pixel-level segmentation, limiting consistent evaluation and applicability. We present SurgMLLMBench, a unified multimodal benchmark explicitly designed for developing and evaluating interactive multimodal LLMs for surgical scene understanding, including the newly collected Micro-surgical Artificial Vascular anastomosIS (MAVIS) dataset. It integrates pixel-level instrument segmentation masks and structured VQA annotations across laparoscopic, robot-assisted, and micro-surgical domains under a unified taxonomy, enabling comprehensive evaluation beyond traditional VQA tasks and richer visual-conversational interactions. Extensive baseline experiments show that a single model trained on SurgMLLMBench achieves consistent performance across domains and generalizes effectively to unseen datasets. SurgMLLMBench will be publicly released as a robust resource to advance multimodal surgical AI research, supporting reproducible evaluation and development of interactive surgical reasoning models.
title SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.21339