SurgPub-Video: A Comprehensive Surgical Video Dataset for Enhanced Surgical Intelligence in Vision-Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yaoqian, Yang, Xikai, Xu, Dunyuan, Yu, Yang, Zhao, Litao, Hu, Xiaowei, Li, Jinpeng, Heng, Pheng-Ann
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912832626360320
author Li, Yaoqian
Yang, Xikai
Xu, Dunyuan
Yu, Yang
Zhao, Litao
Hu, Xiaowei
Li, Jinpeng
Heng, Pheng-Ann
author_facet Li, Yaoqian
Yang, Xikai
Xu, Dunyuan
Yu, Yang
Zhao, Litao
Hu, Xiaowei
Li, Jinpeng
Heng, Pheng-Ann
contents Vision-Language Models (VLMs) have shown significant potential in surgical scene analysis, yet existing models are limited by frame-level datasets and lack high-quality video data with procedural surgical knowledge. To address these challenges, we make the following contributions: (i) SurgPub-Video, a comprehensive dataset of over 3,000 surgical videos and 25 million annotated frames across 11 specialties, sourced from peer-reviewed clinical journals, (ii) SurgLLaVA-Video, a specialized VLM for surgical video understanding, built upon the TinyLLaVA-Video architecture that supports both video-level and frame-level inputs, and (iii) a video-level surgical Visual Question Answering (VQA) benchmark, covering diverse 11 surgical specialities, such as vascular, cardiology, and thoracic. Extensive experiments, conducted on the proposed benchmark and three additional surgical downstream tasks (action recognition, skill assessment, and triplet recognition), show that SurgLLaVA-Video significantly outperforms both general-purpose and surgical-specific VLMs with only three billion parameters. The dataset, model, and benchmark will be released to enable further advancements in surgical video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10054
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SurgPub-Video: A Comprehensive Surgical Video Dataset for Enhanced Surgical Intelligence in Vision-Language Model
Li, Yaoqian
Yang, Xikai
Xu, Dunyuan
Yu, Yang
Zhao, Litao
Hu, Xiaowei
Li, Jinpeng
Heng, Pheng-Ann
Other Quantitative Biology
Vision-Language Models (VLMs) have shown significant potential in surgical scene analysis, yet existing models are limited by frame-level datasets and lack high-quality video data with procedural surgical knowledge. To address these challenges, we make the following contributions: (i) SurgPub-Video, a comprehensive dataset of over 3,000 surgical videos and 25 million annotated frames across 11 specialties, sourced from peer-reviewed clinical journals, (ii) SurgLLaVA-Video, a specialized VLM for surgical video understanding, built upon the TinyLLaVA-Video architecture that supports both video-level and frame-level inputs, and (iii) a video-level surgical Visual Question Answering (VQA) benchmark, covering diverse 11 surgical specialities, such as vascular, cardiology, and thoracic. Extensive experiments, conducted on the proposed benchmark and three additional surgical downstream tasks (action recognition, skill assessment, and triplet recognition), show that SurgLLaVA-Video significantly outperforms both general-purpose and surgical-specific VLMs with only three billion parameters. The dataset, model, and benchmark will be released to enable further advancements in surgical video understanding.
title SurgPub-Video: A Comprehensive Surgical Video Dataset for Enhanced Surgical Intelligence in Vision-Language Model
topic Other Quantitative Biology
url https://arxiv.org/abs/2508.10054