SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Perez, Alejandra, Nwoye, Chinedu, Kermani, Ramtin Raji, Mohareri, Omid, Jamal, Muhammad Abdullah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
by: Perez, Alejandra, et al.
Published: (2026)
by: Perez, Alejandra, et al.
Published: (2026)
Rethinking RGB-D Fusion for Semantic Segmentation in Surgical Datasets
by: Jamal, Muhammad Abdullah, et al.
Published: (2024)
by: Jamal, Muhammad Abdullah, et al.
Published: (2024)
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
by: Han, John J., et al.
Published: (2026)
by: Han, John J., et al.
Published: (2026)
A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders
by: Jamal, Muhammad Abdullah, et al.
Published: (2024)
by: Jamal, Muhammad Abdullah, et al.
Published: (2024)
VidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery
by: Honarmand, Mohammadmahdi, et al.
Published: (2024)
by: Honarmand, Mohammadmahdi, et al.
Published: (2024)
Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition
by: Hamoud, Idris, et al.
Published: (2025)
by: Hamoud, Idris, et al.
Published: (2025)
SurgiTrack: Fine-Grained Multi-Class Multi-Tool Tracking in Surgical Videos
by: Nwoye, Chinedu Innocent, et al.
Published: (2024)
by: Nwoye, Chinedu Innocent, et al.
Published: (2024)
AdaEmbed: Semi-supervised Domain Adaptation in the Embedding Space
by: Mottaghi, Ali, et al.
Published: (2024)
by: Mottaghi, Ali, et al.
Published: (2024)
Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
by: Venkatesh, Danush Kumar, et al.
Published: (2025)
by: Venkatesh, Danush Kumar, et al.
Published: (2025)
Surgical Tattoos in Infrared: A Dataset for Quantifying Tissue Tracking and Mapping
by: Schmidt, Adam, et al.
Published: (2023)
by: Schmidt, Adam, et al.
Published: (2023)
SimGen: A Diffusion-Based Framework for Simultaneous Surgical Image and Segmentation Mask Generation
by: Bhat, Aditya, et al.
Published: (2025)
by: Bhat, Aditya, et al.
Published: (2025)
CholecTrack20: A Multi-Perspective Tracking Dataset for Surgical Tools
by: Nwoye, Chinedu Innocent, et al.
Published: (2023)
by: Nwoye, Chinedu Innocent, et al.
Published: (2023)
Seeing Through Smoke: Surgical Desmoking for Improved Visual Perception
by: Lu, Jingpei, et al.
Published: (2026)
by: Lu, Jingpei, et al.
Published: (2026)
State-Change Learning for Prediction of Future Events in Endoscopic Videos
by: Sharma, Saurav, et al.
Published: (2025)
by: Sharma, Saurav, et al.
Published: (2025)
CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
by: Amangeldi, Aidar, et al.
Published: (2025)
by: Amangeldi, Aidar, et al.
Published: (2025)
Surgical Visual Understanding (SurgVU) Dataset
by: Zia, Aneeq, et al.
Published: (2025)
by: Zia, Aneeq, et al.
Published: (2025)
Surgical Text-to-Image Generation
by: Nwoye, Chinedu Innocent, et al.
Published: (2024)
by: Nwoye, Chinedu Innocent, et al.
Published: (2024)
CoSimGen: Controllable Diffusion Model for Simultaneous Image and Mask Generation
by: Bose, Rupak, et al.
Published: (2025)
by: Bose, Rupak, et al.
Published: (2025)
SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
by: Choi, Tae-Min, et al.
Published: (2025)
by: Choi, Tae-Min, et al.
Published: (2025)
Bridging Vision and Language for Robust Context-Aware Surgical Point Tracking: The VL-SurgPT Dataset and Benchmark
by: Zhou, Rulin, et al.
Published: (2025)
by: Zhou, Rulin, et al.
Published: (2025)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
by: Drago, Mauro Orazio, et al.
Published: (2025)
by: Drago, Mauro Orazio, et al.
Published: (2025)
SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention
by: Dhake, Shreyas C., et al.
Published: (2025)
by: Dhake, Shreyas C., et al.
Published: (2025)
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
by: Zeng, Zhitao, et al.
Published: (2025)
by: Zeng, Zhitao, et al.
Published: (2025)
SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis
by: Wei, Jianhui, et al.
Published: (2025)
by: Wei, Jianhui, et al.
Published: (2025)
Self-supervised Learning via Cluster Distance Prediction for Operating Room Context Awareness
by: Hamoud, Idris, et al.
Published: (2024)
by: Hamoud, Idris, et al.
Published: (2024)
Tracking and Mapping in Medical Computer Vision: A Review
by: Schmidt, Adam, et al.
Published: (2023)
by: Schmidt, Adam, et al.
Published: (2023)
SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?
by: Shin, Jongmin, et al.
Published: (2026)
by: Shin, Jongmin, et al.
Published: (2026)
Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025
by: Zia, Aneeq, et al.
Published: (2023)
by: Zia, Aneeq, et al.
Published: (2023)
Feature Mixing Approach for Detecting Intraoperative Adverse Events in Laparoscopic Roux-en-Y Gastric Bypass Surgery
by: Bose, Rupak, et al.
Published: (2025)
by: Bose, Rupak, et al.
Published: (2025)
SurgFed: Language-guided Multi-Task Federated Learning for Surgical Video Understanding
by: Fang, Zheng, et al.
Published: (2026)
by: Fang, Zheng, et al.
Published: (2026)
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
by: Wang, Guankun, et al.
Published: (2025)
by: Wang, Guankun, et al.
Published: (2025)
SurgPose: a Dataset for Articulated Robotic Surgical Tool Pose Estimation and Tracking
by: Wu, Zijian, et al.
Published: (2025)
by: Wu, Zijian, et al.
Published: (2025)
ViTALS: Vision Transformer for Action Localization in Surgical Nephrectomy
by: Chandra, Soumyadeep, et al.
Published: (2024)
by: Chandra, Soumyadeep, et al.
Published: (2024)
SurgPETL: Parameter-Efficient Image-to-Surgical-Video Transfer Learning for Surgical Phase Recognition
by: Yang, Shu, et al.
Published: (2024)
by: Yang, Shu, et al.
Published: (2024)
Federated Learning for Surgical Vision in Appendicitis Classification: Results of the FedSurg EndoVis 2024 Challenge
by: Kirchner, Max, et al.
Published: (2025)
by: Kirchner, Max, et al.
Published: (2025)
LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning
by: Li, Jiajie, et al.
Published: (2024)
by: Li, Jiajie, et al.
Published: (2024)
SurgPose: Generalisable Surgical Instrument Pose Estimation using Zero-Shot Learning and Stereo Vision
by: Rai, Utsav, et al.
Published: (2025)
by: Rai, Utsav, et al.
Published: (2025)
INTERLACE: Interleaved Layer Pruning and Efficient Adaptation in Large Vision-Language Models
by: Madinei, Parsa, et al.
Published: (2025)
by: Madinei, Parsa, et al.
Published: (2025)
SCARED-C: Corrected Camera Poses for Endoscopic Depth Estimation
by: Han, John J., et al.
Published: (2026)
by: Han, John J., et al.
Published: (2026)
SurgCUT3R: Surgical Scene-Aware Continuous Understanding of Temporal 3D Representation
by: Xu, Kaiyuan, et al.
Published: (2026)
by: Xu, Kaiyuan, et al.
Published: (2026)
Similar Items
-
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
by: Perez, Alejandra, et al.
Published: (2026) -
Rethinking RGB-D Fusion for Semantic Segmentation in Surgical Datasets
by: Jamal, Muhammad Abdullah, et al.
Published: (2024) -
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
by: Han, John J., et al.
Published: (2026) -
A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders
by: Jamal, Muhammad Abdullah, et al.
Published: (2024) -
VidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery
by: Honarmand, Mohammadmahdi, et al.
Published: (2024)