SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Ming, Dong, Wenhui, Zhang, Yang, Zheng, Xiang, Zhang, Zhonghao, Zhou, Zian, Guan, Yunzhi, Xu, Liukun, Peng, Wei, Gong, Zhaoyang, Zhang, Zhicheng, Li, Dachuan, Ma, Xiaosheng, Ma, Yuli, Ni, Jianing, Jiang, Changjiang, Tian, Lixia, Chen, Qixin, Xia, Kaishun, Liu, Pingping, Zhang, Tongshun, Liu, Zhiqiang, Bi, Zhongyan, Si, Chenyang, Sun, Tiansheng, Shan, Caifeng
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911487575982080
author Zhao, Ming
Dong, Wenhui
Zhang, Yang
Zheng, Xiang
Zhang, Zhonghao
Zhou, Zian
Guan, Yunzhi
Xu, Liukun
Peng, Wei
Gong, Zhaoyang
Zhang, Zhicheng
Li, Dachuan
Ma, Xiaosheng
Ma, Yuli
Ni, Jianing
Jiang, Changjiang
Tian, Lixia
Chen, Qixin
Xia, Kaishun
Liu, Pingping
Zhang, Tongshun
Liu, Zhiqiang
Bi, Zhongyan
Si, Chenyang
Sun, Tiansheng
Shan, Caifeng
author_facet Zhao, Ming
Dong, Wenhui
Zhang, Yang
Zheng, Xiang
Zhang, Zhonghao
Zhou, Zian
Guan, Yunzhi
Xu, Liukun
Peng, Wei
Gong, Zhaoyang
Zhang, Zhicheng
Li, Dachuan
Ma, Xiaosheng
Ma, Yuli
Ni, Jianing
Jiang, Changjiang
Tian, Lixia
Chen, Qixin
Xia, Kaishun
Liu, Pingping
Zhang, Tongshun
Liu, Zhiqiang
Bi, Zhongyan
Si, Chenyang
Sun, Tiansheng
Shan, Caifeng
contents Spine disorders affect 619 million people globally and are a leading cause of disability, yet AI-assisted diagnosis remains limited by the lack of level-aware, multimodal datasets. Clinical decision-making for spine disorders requires sophisticated reasoning across X-ray, CT, and MRI at specific vertebral levels. However, progress has been constrained by the absence of traceable, clinically-grounded instruction data and standardized, spine-specific benchmarks. To address this, we introduce SpineMed, an ecosystem co-designed with practicing spine surgeons. It features SpineMed-450k, the first large-scale dataset explicitly designed for vertebral-level reasoning across imaging modalities with over 450,000 instruction instances, and SpineBench, a clinically-grounded evaluation framework. SpineMed-450k is curated from diverse sources, including textbooks, guidelines, open datasets, and ~1,000 de-identified hospital cases, using a clinician-in-the-loop pipeline with a two-stage LLM generation method (draft and revision) to ensure high-quality, traceable data for question-answering, multi-turn consultations, and report generation. SpineBench evaluates models on clinically salient axes, including level identification, pathology assessment, and surgical planning. Our comprehensive evaluation of several recently advanced large vision-language models (LVLMs) on SpineBench reveals systematic weaknesses in fine-grained, level-specific reasoning. In contrast, our model fine-tuned on SpineMed-450k demonstrates consistent and significant improvements across all tasks. Clinician assessments confirm the diagnostic clarity and practical utility of our model's outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus
Zhao, Ming
Dong, Wenhui
Zhang, Yang
Zheng, Xiang
Zhang, Zhonghao
Zhou, Zian
Guan, Yunzhi
Xu, Liukun
Peng, Wei
Gong, Zhaoyang
Zhang, Zhicheng
Li, Dachuan
Ma, Xiaosheng
Ma, Yuli
Ni, Jianing
Jiang, Changjiang
Tian, Lixia
Chen, Qixin
Xia, Kaishun
Liu, Pingping
Zhang, Tongshun
Liu, Zhiqiang
Bi, Zhongyan
Si, Chenyang
Sun, Tiansheng
Shan, Caifeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Spine disorders affect 619 million people globally and are a leading cause of disability, yet AI-assisted diagnosis remains limited by the lack of level-aware, multimodal datasets. Clinical decision-making for spine disorders requires sophisticated reasoning across X-ray, CT, and MRI at specific vertebral levels. However, progress has been constrained by the absence of traceable, clinically-grounded instruction data and standardized, spine-specific benchmarks. To address this, we introduce SpineMed, an ecosystem co-designed with practicing spine surgeons. It features SpineMed-450k, the first large-scale dataset explicitly designed for vertebral-level reasoning across imaging modalities with over 450,000 instruction instances, and SpineBench, a clinically-grounded evaluation framework. SpineMed-450k is curated from diverse sources, including textbooks, guidelines, open datasets, and ~1,000 de-identified hospital cases, using a clinician-in-the-loop pipeline with a two-stage LLM generation method (draft and revision) to ensure high-quality, traceable data for question-answering, multi-turn consultations, and report generation. SpineBench evaluates models on clinically salient axes, including level identification, pathology assessment, and surgical planning. Our comprehensive evaluation of several recently advanced large vision-language models (LVLMs) on SpineBench reveals systematic weaknesses in fine-grained, level-specific reasoning. In contrast, our model fine-tuned on SpineMed-450k demonstrates consistent and significant improvements across all tasks. Clinician assessments confirm the diagnostic clarity and practical utility of our model's outputs.
title SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.03160