Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rau, Anita, Endo, Mark, Aklilu, Josiah, Heo, Jaewoo, Saab, Khaled, Paderno, Alberto, Jopling, Jeffrey, Holsinger, F. Christopher, Yeung-Levy, Serena |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Depth-guided NeRF Training via Earth Mover's Distance
von: Rau, Anita, et al.
Veröffentlicht: (2024)
von: Rau, Anita, et al.
Veröffentlicht: (2024)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
von: Aklilu, Josiah, et al.
Veröffentlicht: (2024)
von: Aklilu, Josiah, et al.
Veröffentlicht: (2024)
Revisiting Active Learning in the Era of Vision Foundation Models
von: Gupte, Sanket Rajan, et al.
Veröffentlicht: (2024)
von: Gupte, Sanket Rajan, et al.
Veröffentlicht: (2024)
Computer Vision Foundation Models in Endoscopy: Proof of Concept in Oropharyngeal Cancer
von: Alberto Paderno, et al.
Veröffentlicht: (2024)
von: Alberto Paderno, et al.
Veröffentlicht: (2024)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
von: Endo, Mark, et al.
Veröffentlicht: (2025)
von: Endo, Mark, et al.
Veröffentlicht: (2025)
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
von: Heo, Jaewoo, et al.
Veröffentlicht: (2024)
von: Heo, Jaewoo, et al.
Veröffentlicht: (2024)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
von: Endo, Mark, et al.
Veröffentlicht: (2024)
von: Endo, Mark, et al.
Veröffentlicht: (2024)
Motion Diffusion-Guided 3D Global HMR from a Dynamic Camera
von: Heo, Jaewoo, et al.
Veröffentlicht: (2024)
von: Heo, Jaewoo, et al.
Veröffentlicht: (2024)
Ask, Pose, Unite: Scaling Data Acquisition for Close Interactions with Vision Language Models
von: Bravo-Sánchez, Laura, et al.
Veröffentlicht: (2024)
von: Bravo-Sánchez, Laura, et al.
Veröffentlicht: (2024)
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
von: Gu, Jeffrey, et al.
Veröffentlicht: (2025)
von: Gu, Jeffrey, et al.
Veröffentlicht: (2025)
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
von: Lozano, Alejandro, et al.
Veröffentlicht: (2025)
von: Lozano, Alejandro, et al.
Veröffentlicht: (2025)
A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
von: Lozano, Alejandro, et al.
Veröffentlicht: (2025)
von: Lozano, Alejandro, et al.
Veröffentlicht: (2025)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
von: Sui, Elaine, et al.
Veröffentlicht: (2024)
von: Sui, Elaine, et al.
Veröffentlicht: (2024)
Video Action Differencing
von: Burgess, James, et al.
Veröffentlicht: (2025)
von: Burgess, James, et al.
Veröffentlicht: (2025)
FIG. 1 in Systematics of the subterranean amphipod genus Bahadzia (Hadziidae), with description of a new species, redescription of B. yagerae, and analysis of phylogeny and biogeography
von: Sawicki, Thomas, et al.
Veröffentlicht: (2004)
von: Sawicki, Thomas, et al.
Veröffentlicht: (2004)
NegVQA: Can Vision Language Models Understand Negation?
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
FIG. 7. Bahadzia yagerae. Paratype female, 5.8 in Systematics of the subterranean amphipod genus Bahadzia (Hadziidae), with description of a new species, redescription of B. yagerae, and analysis of phylogeny and biogeography
von: Sawicki, Thomas, et al.
Veröffentlicht: (2004)
von: Sawicki, Thomas, et al.
Veröffentlicht: (2004)
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
von: Perez, Alejandra, et al.
Veröffentlicht: (2026)
von: Perez, Alejandra, et al.
Veröffentlicht: (2026)
μ-Bench: A Vision-Language Benchmark for Microscopy Understanding
von: Lozano, Alejandro, et al.
Veröffentlicht: (2024)
von: Lozano, Alejandro, et al.
Veröffentlicht: (2024)
Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
Multi-Human Mesh Recovery with Transformers
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
von: Zhang, Yuhui, et al.
Veröffentlicht: (2024)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2024)
Can Large Language Models Match the Conclusions of Systematic Reviews?
von: Polzak, Christopher, et al.
Veröffentlicht: (2025)
von: Polzak, Christopher, et al.
Veröffentlicht: (2025)
Artificial Intelligence in Snoring Sound Analysis: OSA Detection and Obstruction Site Classification, a Systematic Review
von: Francesco Carlo Tartaglia, et al.
Veröffentlicht: (2026)
von: Francesco Carlo Tartaglia, et al.
Veröffentlicht: (2026)
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
von: Letaief, Khaled B., et al.
Veröffentlicht: (2021)
von: Letaief, Khaled B., et al.
Veröffentlicht: (2021)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
CryoHype: Reconstructing a thousand cryo-EM structures with transformer-based hypernetworks
von: Gu, Jeffrey, et al.
Veröffentlicht: (2025)
von: Gu, Jeffrey, et al.
Veröffentlicht: (2025)
Nursing Management of Arteriovenous Fistula and Artificial Intelligence: A Scoping Review
von: Gaetano Ferrara, et al.
Veröffentlicht: (2026)
von: Gaetano Ferrara, et al.
Veröffentlicht: (2026)
Diffusion-HPC: Synthetic Data Generation for Human Mesh Recovery in Challenging Domains
von: Weng, Zhenzhen, et al.
Veröffentlicht: (2023)
von: Weng, Zhenzhen, et al.
Veröffentlicht: (2023)
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
von: Burgess, James, et al.
Veröffentlicht: (2023)
von: Burgess, James, et al.
Veröffentlicht: (2023)
General surgery vision transformer: A video pre-trained foundation model for general surgery
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2024)
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2024)
Artificial Intelligence and Workforce Training
von: Shyamal S. Pandya, et al.
Veröffentlicht: (2026)
von: Shyamal S. Pandya, et al.
Veröffentlicht: (2026)
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
von: Han, John J., et al.
Veröffentlicht: (2026)
von: Han, John J., et al.
Veröffentlicht: (2026)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
Music and Artificial Intelligence: Artistic Trends
von: Pons, Jordi, et al.
Veröffentlicht: (2025)
von: Pons, Jordi, et al.
Veröffentlicht: (2025)
Artificial Intelligence for Women and Child Healthcare: Is AI Able to Change the Beginning of a New Story? A Perspective
von: Patricia Takako Endo
Veröffentlicht: (2025)
von: Patricia Takako Endo
Veröffentlicht: (2025)
Systematic Review: Use of Artificial Intelligence and Unmet Needs in Eosinophilic Oesophagitis
von: Ester Castagnaro, et al.
Veröffentlicht: (2025)
von: Ester Castagnaro, et al.
Veröffentlicht: (2025)
Addressing Water Loss With Artificial Intelligence
von: Jeffrey Johnson
Veröffentlicht: (2025)
von: Jeffrey Johnson
Veröffentlicht: (2025)
Ähnliche Einträge
-
Depth-guided NeRF Training via Earth Mover's Distance
von: Rau, Anita, et al.
Veröffentlicht: (2024) -
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
von: Aklilu, Josiah, et al.
Veröffentlicht: (2024) -
Revisiting Active Learning in the Era of Vision Foundation Models
von: Gupte, Sanket Rajan, et al.
Veröffentlicht: (2024) -
Computer Vision Foundation Models in Endoscopy: Proof of Concept in Oropharyngeal Cancer
von: Alberto Paderno, et al.
Veröffentlicht: (2024) -
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
von: Endo, Mark, et al.
Veröffentlicht: (2025)