Evaluating Large Vision-language Models for Surgical Tool Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Poudel, Nakul, Simon, Richard, Linte, Cristian A. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluation of Intra-operative Patient-specific Methods for Point Cloud Completion for Minimally Invasive Liver Interventions
di: Poudel, Nakul, et al.
Pubblicazione: (2025)
di: Poudel, Nakul, et al.
Pubblicazione: (2025)
Toward Patient-specific Partial Point Cloud to Surface Completion for Pre- to Intra-operative Registration in Image-guided Liver Interventions
di: Poudel, Nakul, et al.
Pubblicazione: (2025)
di: Poudel, Nakul, et al.
Pubblicazione: (2025)
Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence
di: Rau, Anita, et al.
Pubblicazione: (2025)
di: Rau, Anita, et al.
Pubblicazione: (2025)
Assessing Color Vision Test in Large Vision-language Models
di: Ye, Hongfei, et al.
Pubblicazione: (2025)
di: Ye, Hongfei, et al.
Pubblicazione: (2025)
Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models
di: Jin, Juseong, et al.
Pubblicazione: (2024)
di: Jin, Juseong, et al.
Pubblicazione: (2024)
PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures
di: Shukla, Shreya, et al.
Pubblicazione: (2025)
di: Shukla, Shreya, et al.
Pubblicazione: (2025)
Assessing the Performance of the DINOv2 Self-supervised Learning Vision Transformer Model for the Segmentation of the Left Atrium from MRI Images
di: Kundu, Bipasha, et al.
Pubblicazione: (2024)
di: Kundu, Bipasha, et al.
Pubblicazione: (2024)
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
di: Perez, Alejandra, et al.
Pubblicazione: (2026)
di: Perez, Alejandra, et al.
Pubblicazione: (2026)
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
di: Yuan, Kun, et al.
Pubblicazione: (2024)
di: Yuan, Kun, et al.
Pubblicazione: (2024)
Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation
di: Feng, Yongchao, et al.
Pubblicazione: (2025)
di: Feng, Yongchao, et al.
Pubblicazione: (2025)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
di: Seth, Ashish, et al.
Pubblicazione: (2024)
di: Seth, Ashish, et al.
Pubblicazione: (2024)
FlowLearn: Evaluating Large Vision-Language Models on Flowchart Understanding
di: Pan, Huitong, et al.
Pubblicazione: (2024)
di: Pan, Huitong, et al.
Pubblicazione: (2024)
Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models
di: Liang, Shuang, et al.
Pubblicazione: (2025)
di: Liang, Shuang, et al.
Pubblicazione: (2025)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
di: Walimbe, Soham, et al.
Pubblicazione: (2025)
di: Walimbe, Soham, et al.
Pubblicazione: (2025)
StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues
di: Ruan, Zanxi, et al.
Pubblicazione: (2026)
di: Ruan, Zanxi, et al.
Pubblicazione: (2026)
IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models
di: Chen, Zhihao, et al.
Pubblicazione: (2023)
di: Chen, Zhihao, et al.
Pubblicazione: (2023)
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
di: Yan, Bei, et al.
Pubblicazione: (2024)
di: Yan, Bei, et al.
Pubblicazione: (2024)
CholecTrack20: A Multi-Perspective Tracking Dataset for Surgical Tools
di: Nwoye, Chinedu Innocent, et al.
Pubblicazione: (2023)
di: Nwoye, Chinedu Innocent, et al.
Pubblicazione: (2023)
TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos
di: Yu, Yakun, et al.
Pubblicazione: (2026)
di: Yu, Yakun, et al.
Pubblicazione: (2026)
Boundary Constraint-free Biomechanical Model-Based Surface Matching for Intraoperative Liver Deformation Correction
di: Yang, Zixin, et al.
Pubblicazione: (2024)
di: Yang, Zixin, et al.
Pubblicazione: (2024)
Evaluating Large and Lightweight Vision Models for Irregular Component Segmentation in E-Waste Disassembly
di: Zhang, Xinyao, et al.
Pubblicazione: (2026)
di: Zhang, Xinyao, et al.
Pubblicazione: (2026)
Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language Models
di: Babaiee, Zahra, et al.
Pubblicazione: (2025)
di: Babaiee, Zahra, et al.
Pubblicazione: (2025)
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
di: Wang, Sibo, et al.
Pubblicazione: (2024)
di: Wang, Sibo, et al.
Pubblicazione: (2024)
PAS : Prelim Attention Score for Detecting Object Hallucinations in Large Vision--Language Models
di: Hoang-Xuan, Nhat, et al.
Pubblicazione: (2025)
di: Hoang-Xuan, Nhat, et al.
Pubblicazione: (2025)
Leveraging Chat-Based Large Vision Language Models for Multimodal Out-Of-Context Detection
di: Shalabi, Fatma, et al.
Pubblicazione: (2024)
di: Shalabi, Fatma, et al.
Pubblicazione: (2024)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
di: Zhang, Yudong, et al.
Pubblicazione: (2024)
di: Zhang, Yudong, et al.
Pubblicazione: (2024)
EgoSurgery-Tool: A Dataset of Surgical Tool and Hand Detection from Egocentric Open Surgery Videos
di: Fujii, Ryo, et al.
Pubblicazione: (2024)
di: Fujii, Ryo, et al.
Pubblicazione: (2024)
Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI
di: Georgenthum, Hugo, et al.
Pubblicazione: (2025)
di: Georgenthum, Hugo, et al.
Pubblicazione: (2025)
Chain-of-Adaptation: Surgical Vision-Language Adaptation with Reinforcement Learning
di: Li, Jiajie, et al.
Pubblicazione: (2026)
di: Li, Jiajie, et al.
Pubblicazione: (2026)
Auto-Comp: An Automated Pipeline for Scalable Compositional Probing of Contrastive Vision-Language Models
di: Sbrolli, Cristian, et al.
Pubblicazione: (2026)
di: Sbrolli, Cristian, et al.
Pubblicazione: (2026)
Vision language models are unreliable at trivial spatial cognition
di: Khemlani, Sangeet, et al.
Pubblicazione: (2025)
di: Khemlani, Sangeet, et al.
Pubblicazione: (2025)
Evaluating Vision-Language Models for Zero-Shot Detection, Classification, and Association of Motorcycles, Passengers, and Helmets
di: Choi, Lucas, et al.
Pubblicazione: (2024)
di: Choi, Lucas, et al.
Pubblicazione: (2024)
CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models
di: Valente, João, et al.
Pubblicazione: (2025)
di: Valente, João, et al.
Pubblicazione: (2025)
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
di: Shao, Wenqi, et al.
Pubblicazione: (2023)
di: Shao, Wenqi, et al.
Pubblicazione: (2023)
MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context
di: Gu, Zishan, et al.
Pubblicazione: (2024)
di: Gu, Zishan, et al.
Pubblicazione: (2024)
POINTS: Improving Your Vision-language Model with Affordable Strategies
di: Liu, Yuan, et al.
Pubblicazione: (2024)
di: Liu, Yuan, et al.
Pubblicazione: (2024)
Vamos: Versatile Action Models for Video Understanding
di: Wang, Shijie, et al.
Pubblicazione: (2023)
di: Wang, Shijie, et al.
Pubblicazione: (2023)
Representation geometry shapes task performance in vision-language modeling for CT enterography
di: Minoccheri, Cristian, et al.
Pubblicazione: (2026)
di: Minoccheri, Cristian, et al.
Pubblicazione: (2026)
Drone Stereo Vision for Radiata Pine Branch Detection and Distance Measurement: Integrating SGBM and Segmentation Models
di: Lin, Yida, et al.
Pubblicazione: (2024)
di: Lin, Yida, et al.
Pubblicazione: (2024)
SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
di: Xie, Shaoan, et al.
Pubblicazione: (2025)
di: Xie, Shaoan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Evaluation of Intra-operative Patient-specific Methods for Point Cloud Completion for Minimally Invasive Liver Interventions
di: Poudel, Nakul, et al.
Pubblicazione: (2025) -
Toward Patient-specific Partial Point Cloud to Surface Completion for Pre- to Intra-operative Registration in Image-guided Liver Interventions
di: Poudel, Nakul, et al.
Pubblicazione: (2025) -
Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence
di: Rau, Anita, et al.
Pubblicazione: (2025) -
Assessing Color Vision Test in Large Vision-language Models
di: Ye, Hongfei, et al.
Pubblicazione: (2025) -
Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models
di: Jin, Juseong, et al.
Pubblicazione: (2024)