Current validation practice undermines surgical AI development

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Reinke, Annika, Li, Ziying O., Tizabi, Minu D., André, Pascaline, Knopp, Marcel, Rother, Mika M., Machado, Ines P., Altieri, Maria S., Alapatt, Deepak, Bano, Sophia, Bodenstedt, Sebastian, Burgert, Oliver, Chen, Elvis C. S., Collins, Justin W., Colliot, Olivier, Christodoulou, Evangelia, Czempiel, Tobias, Das, Adrito, Docea, Reuben, Donoho, Daniel, Dou, Qi, Eckhoff, Jennifer, Engelhardt, Sandy, Fichtinger, Gabor, Fuernstahl, Philipp, Kilroy, Pablo García, Giannarou, Stamatia, Gilbert, Stephen, Gockel, Ines, Godau, Patrick, Gödeke, Jan, Grantcharov, Teodor P., Haidegger, Tamas, Hann, Alexander, Hashizume, Makoto, Heitz, Charles, Hisey, Rebecca, Hoffmann, Hanna, Huaulmé, Arnaud, Jäger, Paul F., Jannin, Pierre, Jarc, Anthony, Jena, Rohit, Jin, Yueming, Joskowicz, Leo, Joyeux, Luc, Kirchner, Max, Krieger, Axel, Kronreif, Gernot, Lam, Kyle, Laufer, Shlomi, Lavanchy, Joël L., Lee, Gyusung I., Lim, Robert, Liu, Peng, Marcus, Hani J., Mascagni, Pietro, Meireles, Ozanan R., Mueller, Beat P., Mündermann, Lars, Nakawala, Hirenkumar, Navab, Nassir, Ndong, Abdourahmane, Neumann, Juliane, Nickel, Felix, Nolden, Marco, Nwoye, Chinedu, Oh, Namkee, Padoy, Nicolas, Pausch, Thomas, Pfeiffer, Micha, Rädsch, Tim, Ren, Hongliang, Rieke, Nicola, Rivoir, Dominik, Sarikaya, Duygu, Schmidgall, Samuel, Seibold, Matthias, Seidlitz, Silvia, Seitel, Alexander, Sharan, Lalith, Siewerdsen, Jeffrey H., Srivastav, Vinkle, Sznitman, Raphael, Taylor, Russell, Tran, Thuy N., Unberath, Matthias, van der Sommen, Fons, Wagner, Martin, Yamlahi, Amine, Zhou, Shaohua K., Zia, Aneeq, Madani, Amin, Stoyanov, Danail, Speidel, Stefanie, Hashimoto, Daniel A., Kolbinger, Fiona R., Maier-Hein, Lena
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911609836797952
author Reinke, Annika
Li, Ziying O.
Tizabi, Minu D.
André, Pascaline
Knopp, Marcel
Rother, Mika M.
Machado, Ines P.
Altieri, Maria S.
Alapatt, Deepak
Bano, Sophia
Bodenstedt, Sebastian
Burgert, Oliver
Chen, Elvis C. S.
Collins, Justin W.
Colliot, Olivier
Christodoulou, Evangelia
Czempiel, Tobias
Das, Adrito
Docea, Reuben
Donoho, Daniel
Dou, Qi
Eckhoff, Jennifer
Engelhardt, Sandy
Fichtinger, Gabor
Fuernstahl, Philipp
Kilroy, Pablo García
Giannarou, Stamatia
Gilbert, Stephen
Gockel, Ines
Godau, Patrick
Gödeke, Jan
Grantcharov, Teodor P.
Haidegger, Tamas
Hann, Alexander
Hashizume, Makoto
Heitz, Charles
Hisey, Rebecca
Hoffmann, Hanna
Huaulmé, Arnaud
Jäger, Paul F.
Jannin, Pierre
Jarc, Anthony
Jena, Rohit
Jin, Yueming
Joskowicz, Leo
Joyeux, Luc
Kirchner, Max
Krieger, Axel
Kronreif, Gernot
Lam, Kyle
Laufer, Shlomi
Lavanchy, Joël L.
Lee, Gyusung I.
Lim, Robert
Liu, Peng
Marcus, Hani J.
Mascagni, Pietro
Meireles, Ozanan R.
Mueller, Beat P.
Mündermann, Lars
Nakawala, Hirenkumar
Navab, Nassir
Ndong, Abdourahmane
Neumann, Juliane
Nickel, Felix
Nolden, Marco
Nwoye, Chinedu
Oh, Namkee
Padoy, Nicolas
Pausch, Thomas
Pfeiffer, Micha
Rädsch, Tim
Ren, Hongliang
Rieke, Nicola
Rivoir, Dominik
Sarikaya, Duygu
Schmidgall, Samuel
Seibold, Matthias
Seidlitz, Silvia
Seitel, Alexander
Sharan, Lalith
Siewerdsen, Jeffrey H.
Srivastav, Vinkle
Sznitman, Raphael
Taylor, Russell
Tran, Thuy N.
Unberath, Matthias
van der Sommen, Fons
Wagner, Martin
Yamlahi, Amine
Zhou, Shaohua K.
Zia, Aneeq
Madani, Amin
Stoyanov, Danail
Speidel, Stefanie
Hashimoto, Daniel A.
Kolbinger, Fiona R.
Maier-Hein, Lena
author_facet Reinke, Annika
Li, Ziying O.
Tizabi, Minu D.
André, Pascaline
Knopp, Marcel
Rother, Mika M.
Machado, Ines P.
Altieri, Maria S.
Alapatt, Deepak
Bano, Sophia
Bodenstedt, Sebastian
Burgert, Oliver
Chen, Elvis C. S.
Collins, Justin W.
Colliot, Olivier
Christodoulou, Evangelia
Czempiel, Tobias
Das, Adrito
Docea, Reuben
Donoho, Daniel
Dou, Qi
Eckhoff, Jennifer
Engelhardt, Sandy
Fichtinger, Gabor
Fuernstahl, Philipp
Kilroy, Pablo García
Giannarou, Stamatia
Gilbert, Stephen
Gockel, Ines
Godau, Patrick
Gödeke, Jan
Grantcharov, Teodor P.
Haidegger, Tamas
Hann, Alexander
Hashizume, Makoto
Heitz, Charles
Hisey, Rebecca
Hoffmann, Hanna
Huaulmé, Arnaud
Jäger, Paul F.
Jannin, Pierre
Jarc, Anthony
Jena, Rohit
Jin, Yueming
Joskowicz, Leo
Joyeux, Luc
Kirchner, Max
Krieger, Axel
Kronreif, Gernot
Lam, Kyle
Laufer, Shlomi
Lavanchy, Joël L.
Lee, Gyusung I.
Lim, Robert
Liu, Peng
Marcus, Hani J.
Mascagni, Pietro
Meireles, Ozanan R.
Mueller, Beat P.
Mündermann, Lars
Nakawala, Hirenkumar
Navab, Nassir
Ndong, Abdourahmane
Neumann, Juliane
Nickel, Felix
Nolden, Marco
Nwoye, Chinedu
Oh, Namkee
Padoy, Nicolas
Pausch, Thomas
Pfeiffer, Micha
Rädsch, Tim
Ren, Hongliang
Rieke, Nicola
Rivoir, Dominik
Sarikaya, Duygu
Schmidgall, Samuel
Seibold, Matthias
Seidlitz, Silvia
Seitel, Alexander
Sharan, Lalith
Siewerdsen, Jeffrey H.
Srivastav, Vinkle
Sznitman, Raphael
Taylor, Russell
Tran, Thuy N.
Unberath, Matthias
van der Sommen, Fons
Wagner, Martin
Yamlahi, Amine
Zhou, Shaohua K.
Zia, Aneeq
Madani, Amin
Stoyanov, Danail
Speidel, Stefanie
Hashimoto, Daniel A.
Kolbinger, Fiona R.
Maier-Hein, Lena
contents Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation emerging as an important contributing factor. In fact, existing validation practices often neglect the temporal and hierarchical structure of intraoperative videos, producing misleading, unstable, or clinically irrelevant results. In a pioneering, consensus-driven effort, we introduce a comprehensive catalog of validation pitfalls in AI-based surgical video analysis that was derived from a multi-stage Delphi process with 92 international experts. The collected pitfalls span three categories: (1) data (e.g., incomplete annotation, spurious correlations), (2) metric selection and configuration (e.g., neglect of temporal stability, mismatch with clinical needs), and (3) aggregation and reporting (e.g., clinically uninformative aggregation, failure to account for frame dependencies in hierarchical data structures). A systematic review of surgical AI papers reveals that these pitfalls are widespread in current practice, with the majority of studies failing to account for temporal dynamics or hierarchical data structure, or relying on clinically uninformative metrics. Experiments on real surgical video datasets provide empirical evidence that ignoring temporal and hierarchical data structures can substantially understate uncertainty, obscure critical failure modes, and even alter algorithm rankings. To address these shortcomings, we provide a catalogue of best practices compiled in a multi-stage Delphi process. Together, this work provides an evidence-based framework to inform more rigorous validation of surgical video analysis algorithms and to guide future efforts in benchmarking, reporting, regulatory review, and clinical translation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_03769
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Current validation practice undermines surgical AI development
Reinke, Annika
Li, Ziying O.
Tizabi, Minu D.
André, Pascaline
Knopp, Marcel
Rother, Mika M.
Machado, Ines P.
Altieri, Maria S.
Alapatt, Deepak
Bano, Sophia
Bodenstedt, Sebastian
Burgert, Oliver
Chen, Elvis C. S.
Collins, Justin W.
Colliot, Olivier
Christodoulou, Evangelia
Czempiel, Tobias
Das, Adrito
Docea, Reuben
Donoho, Daniel
Dou, Qi
Eckhoff, Jennifer
Engelhardt, Sandy
Fichtinger, Gabor
Fuernstahl, Philipp
Kilroy, Pablo García
Giannarou, Stamatia
Gilbert, Stephen
Gockel, Ines
Godau, Patrick
Gödeke, Jan
Grantcharov, Teodor P.
Haidegger, Tamas
Hann, Alexander
Hashizume, Makoto
Heitz, Charles
Hisey, Rebecca
Hoffmann, Hanna
Huaulmé, Arnaud
Jäger, Paul F.
Jannin, Pierre
Jarc, Anthony
Jena, Rohit
Jin, Yueming
Joskowicz, Leo
Joyeux, Luc
Kirchner, Max
Krieger, Axel
Kronreif, Gernot
Lam, Kyle
Laufer, Shlomi
Lavanchy, Joël L.
Lee, Gyusung I.
Lim, Robert
Liu, Peng
Marcus, Hani J.
Mascagni, Pietro
Meireles, Ozanan R.
Mueller, Beat P.
Mündermann, Lars
Nakawala, Hirenkumar
Navab, Nassir
Ndong, Abdourahmane
Neumann, Juliane
Nickel, Felix
Nolden, Marco
Nwoye, Chinedu
Oh, Namkee
Padoy, Nicolas
Pausch, Thomas
Pfeiffer, Micha
Rädsch, Tim
Ren, Hongliang
Rieke, Nicola
Rivoir, Dominik
Sarikaya, Duygu
Schmidgall, Samuel
Seibold, Matthias
Seidlitz, Silvia
Seitel, Alexander
Sharan, Lalith
Siewerdsen, Jeffrey H.
Srivastav, Vinkle
Sznitman, Raphael
Taylor, Russell
Tran, Thuy N.
Unberath, Matthias
van der Sommen, Fons
Wagner, Martin
Yamlahi, Amine
Zhou, Shaohua K.
Zia, Aneeq
Madani, Amin
Stoyanov, Danail
Speidel, Stefanie
Hashimoto, Daniel A.
Kolbinger, Fiona R.
Maier-Hein, Lena
Other Quantitative Biology
Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation emerging as an important contributing factor. In fact, existing validation practices often neglect the temporal and hierarchical structure of intraoperative videos, producing misleading, unstable, or clinically irrelevant results. In a pioneering, consensus-driven effort, we introduce a comprehensive catalog of validation pitfalls in AI-based surgical video analysis that was derived from a multi-stage Delphi process with 92 international experts. The collected pitfalls span three categories: (1) data (e.g., incomplete annotation, spurious correlations), (2) metric selection and configuration (e.g., neglect of temporal stability, mismatch with clinical needs), and (3) aggregation and reporting (e.g., clinically uninformative aggregation, failure to account for frame dependencies in hierarchical data structures). A systematic review of surgical AI papers reveals that these pitfalls are widespread in current practice, with the majority of studies failing to account for temporal dynamics or hierarchical data structure, or relying on clinically uninformative metrics. Experiments on real surgical video datasets provide empirical evidence that ignoring temporal and hierarchical data structures can substantially understate uncertainty, obscure critical failure modes, and even alter algorithm rankings. To address these shortcomings, we provide a catalogue of best practices compiled in a multi-stage Delphi process. Together, this work provides an evidence-based framework to inform more rigorous validation of surgical video analysis algorithms and to guide future efforts in benchmarking, reporting, regulatory review, and clinical translation.
title Current validation practice undermines surgical AI development
topic Other Quantitative Biology
url https://arxiv.org/abs/2511.03769