Current validation practice undermines surgical AI development
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911609836797952 |
|---|---|
| author | Reinke, Annika Li, Ziying O. Tizabi, Minu D. André, Pascaline Knopp, Marcel Rother, Mika M. Machado, Ines P. Altieri, Maria S. Alapatt, Deepak Bano, Sophia Bodenstedt, Sebastian Burgert, Oliver Chen, Elvis C. S. Collins, Justin W. Colliot, Olivier Christodoulou, Evangelia Czempiel, Tobias Das, Adrito Docea, Reuben Donoho, Daniel Dou, Qi Eckhoff, Jennifer Engelhardt, Sandy Fichtinger, Gabor Fuernstahl, Philipp Kilroy, Pablo García Giannarou, Stamatia Gilbert, Stephen Gockel, Ines Godau, Patrick Gödeke, Jan Grantcharov, Teodor P. Haidegger, Tamas Hann, Alexander Hashizume, Makoto Heitz, Charles Hisey, Rebecca Hoffmann, Hanna Huaulmé, Arnaud Jäger, Paul F. Jannin, Pierre Jarc, Anthony Jena, Rohit Jin, Yueming Joskowicz, Leo Joyeux, Luc Kirchner, Max Krieger, Axel Kronreif, Gernot Lam, Kyle Laufer, Shlomi Lavanchy, Joël L. Lee, Gyusung I. Lim, Robert Liu, Peng Marcus, Hani J. Mascagni, Pietro Meireles, Ozanan R. Mueller, Beat P. Mündermann, Lars Nakawala, Hirenkumar Navab, Nassir Ndong, Abdourahmane Neumann, Juliane Nickel, Felix Nolden, Marco Nwoye, Chinedu Oh, Namkee Padoy, Nicolas Pausch, Thomas Pfeiffer, Micha Rädsch, Tim Ren, Hongliang Rieke, Nicola Rivoir, Dominik Sarikaya, Duygu Schmidgall, Samuel Seibold, Matthias Seidlitz, Silvia Seitel, Alexander Sharan, Lalith Siewerdsen, Jeffrey H. Srivastav, Vinkle Sznitman, Raphael Taylor, Russell Tran, Thuy N. Unberath, Matthias van der Sommen, Fons Wagner, Martin Yamlahi, Amine Zhou, Shaohua K. Zia, Aneeq Madani, Amin Stoyanov, Danail Speidel, Stefanie Hashimoto, Daniel A. Kolbinger, Fiona R. Maier-Hein, Lena |
| author_facet | Reinke, Annika Li, Ziying O. Tizabi, Minu D. André, Pascaline Knopp, Marcel Rother, Mika M. Machado, Ines P. Altieri, Maria S. Alapatt, Deepak Bano, Sophia Bodenstedt, Sebastian Burgert, Oliver Chen, Elvis C. S. Collins, Justin W. Colliot, Olivier Christodoulou, Evangelia Czempiel, Tobias Das, Adrito Docea, Reuben Donoho, Daniel Dou, Qi Eckhoff, Jennifer Engelhardt, Sandy Fichtinger, Gabor Fuernstahl, Philipp Kilroy, Pablo García Giannarou, Stamatia Gilbert, Stephen Gockel, Ines Godau, Patrick Gödeke, Jan Grantcharov, Teodor P. Haidegger, Tamas Hann, Alexander Hashizume, Makoto Heitz, Charles Hisey, Rebecca Hoffmann, Hanna Huaulmé, Arnaud Jäger, Paul F. Jannin, Pierre Jarc, Anthony Jena, Rohit Jin, Yueming Joskowicz, Leo Joyeux, Luc Kirchner, Max Krieger, Axel Kronreif, Gernot Lam, Kyle Laufer, Shlomi Lavanchy, Joël L. Lee, Gyusung I. Lim, Robert Liu, Peng Marcus, Hani J. Mascagni, Pietro Meireles, Ozanan R. Mueller, Beat P. Mündermann, Lars Nakawala, Hirenkumar Navab, Nassir Ndong, Abdourahmane Neumann, Juliane Nickel, Felix Nolden, Marco Nwoye, Chinedu Oh, Namkee Padoy, Nicolas Pausch, Thomas Pfeiffer, Micha Rädsch, Tim Ren, Hongliang Rieke, Nicola Rivoir, Dominik Sarikaya, Duygu Schmidgall, Samuel Seibold, Matthias Seidlitz, Silvia Seitel, Alexander Sharan, Lalith Siewerdsen, Jeffrey H. Srivastav, Vinkle Sznitman, Raphael Taylor, Russell Tran, Thuy N. Unberath, Matthias van der Sommen, Fons Wagner, Martin Yamlahi, Amine Zhou, Shaohua K. Zia, Aneeq Madani, Amin Stoyanov, Danail Speidel, Stefanie Hashimoto, Daniel A. Kolbinger, Fiona R. Maier-Hein, Lena |
| contents | Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation emerging as an important contributing factor. In fact, existing validation practices often neglect the temporal and hierarchical structure of intraoperative videos, producing misleading, unstable, or clinically irrelevant results. In a pioneering, consensus-driven effort, we introduce a comprehensive catalog of validation pitfalls in AI-based surgical video analysis that was derived from a multi-stage Delphi process with 92 international experts. The collected pitfalls span three categories: (1) data (e.g., incomplete annotation, spurious correlations), (2) metric selection and configuration (e.g., neglect of temporal stability, mismatch with clinical needs), and (3) aggregation and reporting (e.g., clinically uninformative aggregation, failure to account for frame dependencies in hierarchical data structures). A systematic review of surgical AI papers reveals that these pitfalls are widespread in current practice, with the majority of studies failing to account for temporal dynamics or hierarchical data structure, or relying on clinically uninformative metrics. Experiments on real surgical video datasets provide empirical evidence that ignoring temporal and hierarchical data structures can substantially understate uncertainty, obscure critical failure modes, and even alter algorithm rankings. To address these shortcomings, we provide a catalogue of best practices compiled in a multi-stage Delphi process. Together, this work provides an evidence-based framework to inform more rigorous validation of surgical video analysis algorithms and to guide future efforts in benchmarking, reporting, regulatory review, and clinical translation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_03769 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Current validation practice undermines surgical AI development Reinke, Annika Li, Ziying O. Tizabi, Minu D. André, Pascaline Knopp, Marcel Rother, Mika M. Machado, Ines P. Altieri, Maria S. Alapatt, Deepak Bano, Sophia Bodenstedt, Sebastian Burgert, Oliver Chen, Elvis C. S. Collins, Justin W. Colliot, Olivier Christodoulou, Evangelia Czempiel, Tobias Das, Adrito Docea, Reuben Donoho, Daniel Dou, Qi Eckhoff, Jennifer Engelhardt, Sandy Fichtinger, Gabor Fuernstahl, Philipp Kilroy, Pablo García Giannarou, Stamatia Gilbert, Stephen Gockel, Ines Godau, Patrick Gödeke, Jan Grantcharov, Teodor P. Haidegger, Tamas Hann, Alexander Hashizume, Makoto Heitz, Charles Hisey, Rebecca Hoffmann, Hanna Huaulmé, Arnaud Jäger, Paul F. Jannin, Pierre Jarc, Anthony Jena, Rohit Jin, Yueming Joskowicz, Leo Joyeux, Luc Kirchner, Max Krieger, Axel Kronreif, Gernot Lam, Kyle Laufer, Shlomi Lavanchy, Joël L. Lee, Gyusung I. Lim, Robert Liu, Peng Marcus, Hani J. Mascagni, Pietro Meireles, Ozanan R. Mueller, Beat P. Mündermann, Lars Nakawala, Hirenkumar Navab, Nassir Ndong, Abdourahmane Neumann, Juliane Nickel, Felix Nolden, Marco Nwoye, Chinedu Oh, Namkee Padoy, Nicolas Pausch, Thomas Pfeiffer, Micha Rädsch, Tim Ren, Hongliang Rieke, Nicola Rivoir, Dominik Sarikaya, Duygu Schmidgall, Samuel Seibold, Matthias Seidlitz, Silvia Seitel, Alexander Sharan, Lalith Siewerdsen, Jeffrey H. Srivastav, Vinkle Sznitman, Raphael Taylor, Russell Tran, Thuy N. Unberath, Matthias van der Sommen, Fons Wagner, Martin Yamlahi, Amine Zhou, Shaohua K. Zia, Aneeq Madani, Amin Stoyanov, Danail Speidel, Stefanie Hashimoto, Daniel A. Kolbinger, Fiona R. Maier-Hein, Lena Other Quantitative Biology Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation emerging as an important contributing factor. In fact, existing validation practices often neglect the temporal and hierarchical structure of intraoperative videos, producing misleading, unstable, or clinically irrelevant results. In a pioneering, consensus-driven effort, we introduce a comprehensive catalog of validation pitfalls in AI-based surgical video analysis that was derived from a multi-stage Delphi process with 92 international experts. The collected pitfalls span three categories: (1) data (e.g., incomplete annotation, spurious correlations), (2) metric selection and configuration (e.g., neglect of temporal stability, mismatch with clinical needs), and (3) aggregation and reporting (e.g., clinically uninformative aggregation, failure to account for frame dependencies in hierarchical data structures). A systematic review of surgical AI papers reveals that these pitfalls are widespread in current practice, with the majority of studies failing to account for temporal dynamics or hierarchical data structure, or relying on clinically uninformative metrics. Experiments on real surgical video datasets provide empirical evidence that ignoring temporal and hierarchical data structures can substantially understate uncertainty, obscure critical failure modes, and even alter algorithm rankings. To address these shortcomings, we provide a catalogue of best practices compiled in a multi-stage Delphi process. Together, this work provides an evidence-based framework to inform more rigorous validation of surgical video analysis algorithms and to guide future efforts in benchmarking, reporting, regulatory review, and clinical translation. |
| title | Current validation practice undermines surgical AI development |
| topic | Other Quantitative Biology |
| url | https://arxiv.org/abs/2511.03769 |