SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Mengya, Huang, Zhongzhen, Imans, Dillan, Ye, Yiru, Zhang, Xiaofan, Dou, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Surgical Action Planning with Large Language Models
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
Generalized Recognition of Basic Surgical Actions Enables Skill Assessment and Vision-Language-Model-based Surgical Planning
von: Xu, Mengya, et al.
Veröffentlicht: (2026)
von: Xu, Mengya, et al.
Veröffentlicht: (2026)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
von: Chen, Yi, et al.
Veröffentlicht: (2023)
von: Chen, Yi, et al.
Veröffentlicht: (2023)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling
von: He, Yufan, et al.
Veröffentlicht: (2025)
von: He, Yufan, et al.
Veröffentlicht: (2025)
Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
Clinical Graph-Mediated Distillation for Unpaired MRI-to-CFI Hypertension Prediction
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
Unsupervised Domain Adaptation with SAM-RefiSeR for Enhanced Brain Tumor Segmentation
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
von: Du, Yiyang, et al.
Veröffentlicht: (2025)
von: Du, Yiyang, et al.
Veröffentlicht: (2025)
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
von: Chen, Liang, et al.
Veröffentlicht: (2024)
von: Chen, Liang, et al.
Veröffentlicht: (2024)
Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
A Survey on Benchmarks of Multimodal Large Language Models
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
UCAgents: Unidirectional Convergence for Visual Evidence Anchored Multi-Agent Medical Decision-Making
von: Feng, Qianhan, et al.
Veröffentlicht: (2025)
von: Feng, Qianhan, et al.
Veröffentlicht: (2025)
LLAVADI: What Matters For Multimodal Large Language Models Distillation
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
Frequency-Modulated Visual Restoration for Matryoshka Large Multimodal Models
von: Pan, Qingtao, et al.
Veröffentlicht: (2026)
von: Pan, Qingtao, et al.
Veröffentlicht: (2026)
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2023)
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2023)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
von: Niu, Tianyi, et al.
Veröffentlicht: (2025)
von: Niu, Tianyi, et al.
Veröffentlicht: (2025)
Benchmarking Multimodal Large Language Models for Face Recognition
von: Shahreza, Hatef Otroshi, et al.
Veröffentlicht: (2025)
von: Shahreza, Hatef Otroshi, et al.
Veröffentlicht: (2025)
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
Benchmarking Large Language Models for Image Classification of Marine Mammals
von: Qi, Yijiashun, et al.
Veröffentlicht: (2024)
von: Qi, Yijiashun, et al.
Veröffentlicht: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks
von: Hong, Jindong, et al.
Veröffentlicht: (2025)
von: Hong, Jindong, et al.
Veröffentlicht: (2025)
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
von: Kang, Zhaolu, et al.
Veröffentlicht: (2025)
von: Kang, Zhaolu, et al.
Veröffentlicht: (2025)
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Surgical Action Planning with Large Language Models
von: Xu, Mengya, et al.
Veröffentlicht: (2025) -
Generalized Recognition of Basic Surgical Actions Enables Skill Assessment and Vision-Language-Model-based Surgical Planning
von: Xu, Mengya, et al.
Veröffentlicht: (2026) -
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
von: Chen, Yi, et al.
Veröffentlicht: (2023) -
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024) -
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)