Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Rau, Anita, Endo, Mark, Aklilu, Josiah, Heo, Jaewoo, Saab, Khaled, Paderno, Alberto, Jopling, Jeffrey, Holsinger, F. Christopher, Yeung-Levy, Serena |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Depth-guided NeRF Training via Earth Mover's Distance
by: Rau, Anita, et al.
Published: (2024)
by: Rau, Anita, et al.
Published: (2024)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
by: Aklilu, Josiah, et al.
Published: (2024)
by: Aklilu, Josiah, et al.
Published: (2024)
Revisiting Active Learning in the Era of Vision Foundation Models
by: Gupte, Sanket Rajan, et al.
Published: (2024)
by: Gupte, Sanket Rajan, et al.
Published: (2024)
Computer Vision Foundation Models in Endoscopy: Proof of Concept in Oropharyngeal Cancer
by: Alberto Paderno, et al.
Published: (2024)
by: Alberto Paderno, et al.
Published: (2024)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
by: Endo, Mark, et al.
Published: (2025)
by: Endo, Mark, et al.
Published: (2025)
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
by: Heo, Jaewoo, et al.
Published: (2024)
by: Heo, Jaewoo, et al.
Published: (2024)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
by: Endo, Mark, et al.
Published: (2024)
by: Endo, Mark, et al.
Published: (2024)
Motion Diffusion-Guided 3D Global HMR from a Dynamic Camera
by: Heo, Jaewoo, et al.
Published: (2024)
by: Heo, Jaewoo, et al.
Published: (2024)
Ask, Pose, Unite: Scaling Data Acquisition for Close Interactions with Vision Language Models
by: Bravo-Sánchez, Laura, et al.
Published: (2024)
by: Bravo-Sánchez, Laura, et al.
Published: (2024)
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
by: Zhang, Yuhui, et al.
Published: (2025)
by: Zhang, Yuhui, et al.
Published: (2025)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
by: Gu, Jeffrey, et al.
Published: (2025)
by: Gu, Jeffrey, et al.
Published: (2025)
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
by: Lozano, Alejandro, et al.
Published: (2025)
by: Lozano, Alejandro, et al.
Published: (2025)
A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
by: Lozano, Alejandro, et al.
Published: (2025)
by: Lozano, Alejandro, et al.
Published: (2025)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
by: Sui, Elaine, et al.
Published: (2024)
by: Sui, Elaine, et al.
Published: (2024)
Video Action Differencing
by: Burgess, James, et al.
Published: (2025)
by: Burgess, James, et al.
Published: (2025)
FIG. 1 in Systematics of the subterranean amphipod genus Bahadzia (Hadziidae), with description of a new species, redescription of B. yagerae, and analysis of phylogeny and biogeography
by: Sawicki, Thomas, et al.
Published: (2004)
by: Sawicki, Thomas, et al.
Published: (2004)
NegVQA: Can Vision Language Models Understand Negation?
by: Zhang, Yuhui, et al.
Published: (2025)
by: Zhang, Yuhui, et al.
Published: (2025)
FIG. 7. Bahadzia yagerae. Paratype female, 5.8 in Systematics of the subterranean amphipod genus Bahadzia (Hadziidae), with description of a new species, redescription of B. yagerae, and analysis of phylogeny and biogeography
by: Sawicki, Thomas, et al.
Published: (2004)
by: Sawicki, Thomas, et al.
Published: (2004)
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
by: Perez, Alejandra, et al.
Published: (2026)
by: Perez, Alejandra, et al.
Published: (2026)
μ-Bench: A Vision-Language Benchmark for Microscopy Understanding
by: Lozano, Alejandro, et al.
Published: (2024)
by: Lozano, Alejandro, et al.
Published: (2024)
Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
Multi-Human Mesh Recovery with Transformers
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
Can Large Language Models Match the Conclusions of Systematic Reviews?
by: Polzak, Christopher, et al.
Published: (2025)
by: Polzak, Christopher, et al.
Published: (2025)
Artificial Intelligence in Snoring Sound Analysis: OSA Detection and Obstruction Site Classification, a Systematic Review
by: Francesco Carlo Tartaglia, et al.
Published: (2026)
by: Francesco Carlo Tartaglia, et al.
Published: (2026)
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
by: Zeng, Zhitao, et al.
Published: (2025)
by: Zeng, Zhitao, et al.
Published: (2025)
Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
by: Letaief, Khaled B., et al.
Published: (2021)
by: Letaief, Khaled B., et al.
Published: (2021)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
by: Chen, Liangyu, et al.
Published: (2025)
by: Chen, Liangyu, et al.
Published: (2025)
CryoHype: Reconstructing a thousand cryo-EM structures with transformer-based hypernetworks
by: Gu, Jeffrey, et al.
Published: (2025)
by: Gu, Jeffrey, et al.
Published: (2025)
Nursing Management of Arteriovenous Fistula and Artificial Intelligence: A Scoping Review
by: Gaetano Ferrara, et al.
Published: (2026)
by: Gaetano Ferrara, et al.
Published: (2026)
Diffusion-HPC: Synthetic Data Generation for Human Mesh Recovery in Challenging Domains
by: Weng, Zhenzhen, et al.
Published: (2023)
by: Weng, Zhenzhen, et al.
Published: (2023)
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
by: Burgess, James, et al.
Published: (2023)
by: Burgess, James, et al.
Published: (2023)
General surgery vision transformer: A video pre-trained foundation model for general surgery
by: Schmidgall, Samuel, et al.
Published: (2024)
by: Schmidgall, Samuel, et al.
Published: (2024)
Artificial Intelligence and Workforce Training
by: Shyamal S. Pandya, et al.
Published: (2026)
by: Shyamal S. Pandya, et al.
Published: (2026)
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
by: Han, John J., et al.
Published: (2026)
by: Han, John J., et al.
Published: (2026)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
Music and Artificial Intelligence: Artistic Trends
by: Pons, Jordi, et al.
Published: (2025)
by: Pons, Jordi, et al.
Published: (2025)
Artificial Intelligence for Women and Child Healthcare: Is AI Able to Change the Beginning of a New Story? A Perspective
by: Patricia Takako Endo
Published: (2025)
by: Patricia Takako Endo
Published: (2025)
Systematic Review: Use of Artificial Intelligence and Unmet Needs in Eosinophilic Oesophagitis
by: Ester Castagnaro, et al.
Published: (2025)
by: Ester Castagnaro, et al.
Published: (2025)
Addressing Water Loss With Artificial Intelligence
by: Jeffrey Johnson
Published: (2025)
by: Jeffrey Johnson
Published: (2025)
Similar Items
-
Depth-guided NeRF Training via Earth Mover's Distance
by: Rau, Anita, et al.
Published: (2024) -
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
by: Aklilu, Josiah, et al.
Published: (2024) -
Revisiting Active Learning in the Era of Vision Foundation Models
by: Gupte, Sanket Rajan, et al.
Published: (2024) -
Computer Vision Foundation Models in Endoscopy: Proof of Concept in Oropharyngeal Cancer
by: Alberto Paderno, et al.
Published: (2024) -
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
by: Endo, Mark, et al.
Published: (2025)