Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study
Fuente:
arXiv
Saved in:
| Main Authors: | Mayer, Leon, Rädsch, Tim, Michael, Dominik, Luttner, Lucas, Yamlahi, Amine, Christodoulou, Evangelia, Godau, Patrick, Knopp, Marcel, Reinke, Annika, Kolbinger, Fiona, Maier-Hein, Lena |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models
by: Mayer, Leon, et al.
Published: (2025)
by: Mayer, Leon, et al.
Published: (2025)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
by: Rädsch, Tim, et al.
Published: (2025)
by: Rädsch, Tim, et al.
Published: (2025)
Advancing standards in biomedical image analysis validation: A perspective on Metrics Reloaded
by: Annika Reinke, et al.
Published: (2025)
by: Annika Reinke, et al.
Published: (2025)
Quality Assured: Rethinking Annotation Strategies in Imaging AI
by: Rädsch, Tim, et al.
Published: (2024)
by: Rädsch, Tim, et al.
Published: (2024)
Beyond Knowledge Silos: Task Fingerprinting for Democratization of Medical Imaging AI
by: Godau, Patrick, et al.
Published: (2024)
by: Godau, Patrick, et al.
Published: (2024)
Training of Neural Networks with Uncertain Data: A Mixture of Experts Approach
by: Luttner, Lucas
Published: (2023)
by: Luttner, Lucas
Published: (2023)
Performance uncertainty in medical image analysis: a large-scale investigation of confidence intervals
by: André, Pascaline, et al.
Published: (2026)
by: André, Pascaline, et al.
Published: (2026)
Data Augmentation for Surgical Scene Segmentation with Anatomy-Aware Diffusion Models
by: Venkatesh, Danush Kumar, et al.
Published: (2024)
by: Venkatesh, Danush Kumar, et al.
Published: (2024)
Federated Learning for Surgical Vision in Appendicitis Classification: Results of the FedSurg EndoVis 2024 Challenge
by: Kirchner, Max, et al.
Published: (2025)
by: Kirchner, Max, et al.
Published: (2025)
Exploring Semantic Consistency in Unpaired Image Translation to Generate Data for Surgical Applications
by: Venkatesh, Danush Kumar, et al.
Published: (2023)
by: Venkatesh, Danush Kumar, et al.
Published: (2023)
FISBe: A real-world benchmark dataset for instance segmentation of long-range thin filamentous structures
by: Mais, Lisa, et al.
Published: (2024)
by: Mais, Lisa, et al.
Published: (2024)
False Promises in Medical Imaging AI? Assessing Validity of Outperformance Claims
by: Christodoulou, Evangelia, et al.
Published: (2025)
by: Christodoulou, Evangelia, et al.
Published: (2025)
Scientific approaches to technological officiating aids in game sports
by: Otto Kolbinger
Published: (2017)
by: Otto Kolbinger
Published: (2017)
Application-driven Validation of Posteriors in Inverse Problems
by: Adler, Tim J., et al.
Published: (2023)
by: Adler, Tim J., et al.
Published: (2023)
Strategies to Improve Real-World Applicability of Laparoscopic Anatomy Segmentation Models
by: Kolbinger, Fiona R., et al.
Published: (2024)
by: Kolbinger, Fiona R., et al.
Published: (2024)
Confidence intervals uncovered: Are we ready for real-world medical imaging AI?
by: Christodoulou, Evangelia, et al.
Published: (2024)
by: Christodoulou, Evangelia, et al.
Published: (2024)
Unsupervised Latent Stain Adaptation for Computational Pathology
by: Reisenbüchler, Daniel, et al.
Published: (2024)
by: Reisenbüchler, Daniel, et al.
Published: (2024)
Field strength-dependent performance variability in deep learning-based analysis of magnetic resonance imaging
by: Qadir, Muhammad Ibtsaam, et al.
Published: (2025)
by: Qadir, Muhammad Ibtsaam, et al.
Published: (2025)
The Influence of Width Ratios on Structural Beauty in Male Faces
by: Tennstedt, Theresa, et al.
Published: (2026)
by: Tennstedt, Theresa, et al.
Published: (2026)
Adaptive-CaRe: Adaptive Causal Regularization for Robust Outcome Prediction
by: Bhasker, Nithya, et al.
Published: (2026)
by: Bhasker, Nithya, et al.
Published: (2026)
Handling Geometric Domain Shifts in Semantic Segmentation of Surgical RGB and Hyperspectral Images
by: Seidlitz, Silvia, et al.
Published: (2024)
by: Seidlitz, Silvia, et al.
Published: (2024)
Continual Domain Incremental Learning for Privacy-aware Digital Pathology
by: Kumari, Pratibha, et al.
Published: (2024)
by: Kumari, Pratibha, et al.
Published: (2024)
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
by: Kahl, Kim-Celine, et al.
Published: (2024)
by: Kahl, Kim-Celine, et al.
Published: (2024)
Divide and Conquer: A Large-Scale Dataset and Model for Left-Right Breast MRI Segmentation
by: Rokuss, Maximilian, et al.
Published: (2025)
by: Rokuss, Maximilian, et al.
Published: (2025)
Overcoming Common Flaws in the Evaluation of Selective Classification Systems
by: Traub, Jeremias, et al.
Published: (2024)
by: Traub, Jeremias, et al.
Published: (2024)
AI‐driven preoperative risk assessment in kidney cancer surgery: A comparative feasibility study of machine learning models
by: Julia Mühlbauer, et al.
Published: (2025)
by: Julia Mühlbauer, et al.
Published: (2025)
Praxisform Polizeinotruf
by: Knopp, Philipp
Published: (2026)
by: Knopp, Philipp
Published: (2026)
Theory of functions / Konrad Knopp ; translated by Frederick Bagemihl
by: Knopp, Konrad
by: Knopp, Konrad
Station list and links to master tracks in different resolutions of POLARSTERN cruise ANT-XVI/4, Cape Town - Bremerhaven, 1999-05-11 - 1999-06-03
by: Reinke, Manfred
Published: (2015)
by: Reinke, Manfred
Published: (2015)
Station list and links to master tracks in different resolutions of POLARSTERN cruise ANT-XIX/1, Bremerhaven - Cape Town, 2001-11-07 - 2001-11-30
by: Reinke, Manfred
Published: (2015)
by: Reinke, Manfred
Published: (2015)
One model to use them all: Training a segmentation model with complementary datasets
by: Jenke, Alexander C., et al.
Published: (2024)
by: Jenke, Alexander C., et al.
Published: (2024)
Learned Discrepancy Reconstruction and Benchmark Dataset for Magnetic Particle Imaging
by: Iske, Meira, et al.
Published: (2025)
by: Iske, Meira, et al.
Published: (2025)
Reading Decisions from Gaze Direction during Graphics Turing Test of Gait Animation
by: Knopp, Benjamin, et al.
Published: (2025)
by: Knopp, Benjamin, et al.
Published: (2025)
Preference Redirection via Attention Concentration: An Attack on Computer Use Agents
by: Seip, Dominik, et al.
Published: (2026)
by: Seip, Dominik, et al.
Published: (2026)
The Federated Tumor Segmentation (FeTS) Challenge
by: Pati, Sarthak, et al.
Published: (2021)
by: Pati, Sarthak, et al.
Published: (2021)
Toward Scalable Automated Repository-Level Datasets for Software Vulnerability Detection
by: Lbath, Amine
Published: (2026)
by: Lbath, Amine
Published: (2026)
Super Phantoms: advanced models for testing medical imaging technologies
by: Manohar, Srirang, et al.
Published: (2023)
by: Manohar, Srirang, et al.
Published: (2023)
Surgical Phase and Instrument Recognition: How to identify appropriate Dataset Splits
by: Kostiuchik, Georgii, et al.
Published: (2023)
by: Kostiuchik, Georgii, et al.
Published: (2023)
Automatic Calibration of a Multi-Camera System with Limited Overlapping Fields of View for 3D Surgical Scene Reconstruction
by: Flückiger, Tim, et al.
Published: (2025)
by: Flückiger, Tim, et al.
Published: (2025)
A multi-center analysis of deep learning methods for video polyp detection and segmentation
by: Ghatwary, Noha, et al.
Published: (2026)
by: Ghatwary, Noha, et al.
Published: (2026)
Similar Items
-
6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models
by: Mayer, Leon, et al.
Published: (2025) -
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
by: Rädsch, Tim, et al.
Published: (2025) -
Advancing standards in biomedical image analysis validation: A perspective on Metrics Reloaded
by: Annika Reinke, et al.
Published: (2025) -
Quality Assured: Rethinking Annotation Strategies in Imaging AI
by: Rädsch, Tim, et al.
Published: (2024) -
Beyond Knowledge Silos: Task Fingerprinting for Democratization of Medical Imaging AI
by: Godau, Patrick, et al.
Published: (2024)