Data-Efficient Surgical Phase Segmentation in Small-Incision Cataract Surgery: A Controlled Study of Vision Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Spencer, Lincoln, Wang, Song, Chen, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Phase-Informed Tool Segmentation for Manual Small-Incision Cataract Surgery
by: Sachdeva, Bhuvan, et al.
Published: (2024)
by: Sachdeva, Bhuvan, et al.
Published: (2024)
Surgical-DINO: Adapter Learning of Foundation Models for Depth Estimation in Endoscopic Surgery
by: Cui, Beilei, et al.
Published: (2024)
by: Cui, Beilei, et al.
Published: (2024)
EndoARSS: Adapting Spatially-Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery
by: Wang, Guankun, et al.
Published: (2025)
by: Wang, Guankun, et al.
Published: (2025)
Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
by: Li, Qiankun, et al.
Published: (2025)
by: Li, Qiankun, et al.
Published: (2025)
Small Object Few-shot Segmentation for Vision-based Industrial Inspection
by: Zhang, Zilong, et al.
Published: (2024)
by: Zhang, Zilong, et al.
Published: (2024)
CataractSAM-2: A Domain-Adapted Model for Anterior Segment Surgery Segmentation and Scalable Ground-Truth Annotation
by: Eslami, Mohammad, et al.
Published: (2026)
by: Eslami, Mohammad, et al.
Published: (2026)
Reprogramming Vision Foundation Models for Spatio-Temporal Forecasting
by: Chen, Changlu, et al.
Published: (2025)
by: Chen, Changlu, et al.
Published: (2025)
Boosting Medical Image Classification with Segmentation Foundation Model
by: Gu, Pengfei, et al.
Published: (2024)
by: Gu, Pengfei, et al.
Published: (2024)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
by: Deng, Andong, et al.
Published: (2025)
by: Deng, Andong, et al.
Published: (2025)
Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis
by: Chen, Tingxuan, et al.
Published: (2025)
by: Chen, Tingxuan, et al.
Published: (2025)
Label-Efficient Data Augmentation with Video Diffusion Models for Guidewire Segmentation in Cardiac Fluoroscopy
by: Pan, Shaoyan, et al.
Published: (2024)
by: Pan, Shaoyan, et al.
Published: (2024)
Topology-Driven Transferability Estimation of Medical Foundation Models for Segmentation
by: Tang, Jiaqi, et al.
Published: (2026)
by: Tang, Jiaqi, et al.
Published: (2026)
EgoSurgery-Phase: A Dataset of Surgical Phase Recognition from Egocentric Open Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
VLM-Guided Iterative Refinement for Surgical Image Segmentation with Foundation Models
by: Lou, Ange, et al.
Published: (2026)
by: Lou, Ange, et al.
Published: (2026)
Stabilizing Temporal Inference Dynamics for Online Surgical Phase Recognition
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Robust SAM: On the Adversarial Robustness of Vision Foundation Models
by: Long, Jiahuan, et al.
Published: (2025)
by: Long, Jiahuan, et al.
Published: (2025)
LKM-UNet: Large Kernel Vision Mamba UNet for Medical Image Segmentation
by: Wang, Jinhong, et al.
Published: (2024)
by: Wang, Jinhong, et al.
Published: (2024)
Intelligent Communication Mixture-of-Experts Boosted-Medical Image Segmentation Foundation Model
by: Zhang, Xinwei, et al.
Published: (2025)
by: Zhang, Xinwei, et al.
Published: (2025)
Playing to Vision Foundation Model's Strengths in Stereo Matching
by: Liu, Chuang-Wei, et al.
Published: (2024)
by: Liu, Chuang-Wei, et al.
Published: (2024)
Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models
by: Jin, Juseong, et al.
Published: (2024)
by: Jin, Juseong, et al.
Published: (2024)
CAT-SG: A Large Dynamic Scene Graph Dataset for Fine-Grained Understanding of Cataract Surgery
by: Holm, Felix, et al.
Published: (2025)
by: Holm, Felix, et al.
Published: (2025)
Surgical-MambaLLM: Mamba2-enhanced Multimodal Large Language Model for VQLA in Robotic Surgery
by: Hao, Pengfei, et al.
Published: (2025)
by: Hao, Pengfei, et al.
Published: (2025)
Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
by: Xuan, Yunyi, et al.
Published: (2024)
by: Xuan, Yunyi, et al.
Published: (2024)
Understanding the Transfer Limits of Vision Foundation Models
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation
by: Zhu, Yuanbing, et al.
Published: (2024)
by: Zhu, Yuanbing, et al.
Published: (2024)
Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis
by: Ahmadi, Mohammad Javad, et al.
Published: (2025)
by: Ahmadi, Mohammad Javad, et al.
Published: (2025)
DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models
by: Sufian, Abu, et al.
Published: (2025)
by: Sufian, Abu, et al.
Published: (2025)
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
by: Perez, Alejandra, et al.
Published: (2026)
by: Perez, Alejandra, et al.
Published: (2026)
Evaluating Large Vision-language Models for Surgical Tool Detection
by: Poudel, Nakul, et al.
Published: (2026)
by: Poudel, Nakul, et al.
Published: (2026)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning
by: Hao, Pengfei, et al.
Published: (2025)
by: Hao, Pengfei, et al.
Published: (2025)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
by: Deng, Andong, et al.
Published: (2024)
by: Deng, Andong, et al.
Published: (2024)
Holistic Surgical Phase Recognition with Hierarchical Input Dependent State Space Models
by: Wu, Haoyang, et al.
Published: (2025)
by: Wu, Haoyang, et al.
Published: (2025)
PLPHP: Per-Layer Per-Head Vision Token Pruning for Efficient Large Vision-Language Models
by: Meng, Yu, et al.
Published: (2025)
by: Meng, Yu, et al.
Published: (2025)
The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery
by: Wang, Guankun, et al.
Published: (2024)
by: Wang, Guankun, et al.
Published: (2024)
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model
by: Chen, Yi-Chia, et al.
Published: (2024)
by: Chen, Yi-Chia, et al.
Published: (2024)
Medical SAM3: A Foundation Model for Universal Prompt-Driven Medical Image Segmentation
by: Jiang, Chongcong, et al.
Published: (2026)
by: Jiang, Chongcong, et al.
Published: (2026)
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
by: He, Jialuo, et al.
Published: (2026)
by: He, Jialuo, et al.
Published: (2026)
Towards Dynamic and Small Objects Refinement for Unsupervised Domain Adaptative Nighttime Semantic Segmentation
by: Pan, Jingyi, et al.
Published: (2023)
by: Pan, Jingyi, et al.
Published: (2023)
Similar Items
-
Phase-Informed Tool Segmentation for Manual Small-Incision Cataract Surgery
by: Sachdeva, Bhuvan, et al.
Published: (2024) -
Surgical-DINO: Adapter Learning of Foundation Models for Depth Estimation in Endoscopic Surgery
by: Cui, Beilei, et al.
Published: (2024) -
EndoARSS: Adapting Spatially-Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery
by: Wang, Guankun, et al.
Published: (2025) -
Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
by: Li, Qiankun, et al.
Published: (2025) -
Small Object Few-shot Segmentation for Vision-based Industrial Inspection
by: Zhang, Zilong, et al.
Published: (2024)