On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, John J., Schmidt, Adam, Jamal, Muhammad Abdullah, Nwoye, Chinedu, Rau, Anita, Wu, Jie Ying, Mohareri, Omid |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
di: Perez, Alejandra, et al.
Pubblicazione: (2026)
di: Perez, Alejandra, et al.
Pubblicazione: (2026)
Rethinking RGB-D Fusion for Semantic Segmentation in Surgical Datasets
di: Jamal, Muhammad Abdullah, et al.
Pubblicazione: (2024)
di: Jamal, Muhammad Abdullah, et al.
Pubblicazione: (2024)
SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
di: Perez, Alejandra, et al.
Pubblicazione: (2025)
di: Perez, Alejandra, et al.
Pubblicazione: (2025)
A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders
di: Jamal, Muhammad Abdullah, et al.
Pubblicazione: (2024)
di: Jamal, Muhammad Abdullah, et al.
Pubblicazione: (2024)
Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
di: Venkatesh, Danush Kumar, et al.
Pubblicazione: (2025)
di: Venkatesh, Danush Kumar, et al.
Pubblicazione: (2025)
VidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery
di: Honarmand, Mohammadmahdi, et al.
Pubblicazione: (2024)
di: Honarmand, Mohammadmahdi, et al.
Pubblicazione: (2024)
SCARED-C: Corrected Camera Poses for Endoscopic Depth Estimation
di: Han, John J., et al.
Pubblicazione: (2026)
di: Han, John J., et al.
Pubblicazione: (2026)
Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition
di: Hamoud, Idris, et al.
Pubblicazione: (2025)
di: Hamoud, Idris, et al.
Pubblicazione: (2025)
SurgiTrack: Fine-Grained Multi-Class Multi-Tool Tracking in Surgical Videos
di: Nwoye, Chinedu Innocent, et al.
Pubblicazione: (2024)
di: Nwoye, Chinedu Innocent, et al.
Pubblicazione: (2024)
Surgical Tattoos in Infrared: A Dataset for Quantifying Tissue Tracking and Mapping
di: Schmidt, Adam, et al.
Pubblicazione: (2023)
di: Schmidt, Adam, et al.
Pubblicazione: (2023)
AdaEmbed: Semi-supervised Domain Adaptation in the Embedding Space
di: Mottaghi, Ali, et al.
Pubblicazione: (2024)
di: Mottaghi, Ali, et al.
Pubblicazione: (2024)
SimGen: A Diffusion-Based Framework for Simultaneous Surgical Image and Segmentation Mask Generation
di: Bhat, Aditya, et al.
Pubblicazione: (2025)
di: Bhat, Aditya, et al.
Pubblicazione: (2025)
Tracking and Mapping in Medical Computer Vision: A Review
di: Schmidt, Adam, et al.
Pubblicazione: (2023)
di: Schmidt, Adam, et al.
Pubblicazione: (2023)
Seeing Through Smoke: Surgical Desmoking for Improved Visual Perception
di: Lu, Jingpei, et al.
Pubblicazione: (2026)
di: Lu, Jingpei, et al.
Pubblicazione: (2026)
CholecTrack20: A Multi-Perspective Tracking Dataset for Surgical Tools
di: Nwoye, Chinedu Innocent, et al.
Pubblicazione: (2023)
di: Nwoye, Chinedu Innocent, et al.
Pubblicazione: (2023)
Depth Helps: Improving Pre-trained RGB-based Policy with Depth Information Injection
di: Pang, Xincheng, et al.
Pubblicazione: (2024)
di: Pang, Xincheng, et al.
Pubblicazione: (2024)
Surgical Text-to-Image Generation
di: Nwoye, Chinedu Innocent, et al.
Pubblicazione: (2024)
di: Nwoye, Chinedu Innocent, et al.
Pubblicazione: (2024)
CoSimGen: Controllable Diffusion Model for Simultaneous Image and Mask Generation
di: Bose, Rupak, et al.
Pubblicazione: (2025)
di: Bose, Rupak, et al.
Pubblicazione: (2025)
State-Change Learning for Prediction of Future Events in Endoscopic Videos
di: Sharma, Saurav, et al.
Pubblicazione: (2025)
di: Sharma, Saurav, et al.
Pubblicazione: (2025)
The Consequences of Extending a Country's Library Legislation to the Inclusion of Academic Libraries, with Special Reference to Nigeria.
di: Nwoye, S. C.
Pubblicazione: (1979)
di: Nwoye, S. C.
Pubblicazione: (1979)
The distress of one‐dimensional fertility in an African family
di: Augustine Nwoye
Pubblicazione: (2024)
di: Augustine Nwoye
Pubblicazione: (2024)
Predicting Depth Maps from Single RGB Images and Addressing Missing Information in Depth Estimation
di: Chaar, Mohamad Mofeed, et al.
Pubblicazione: (2025)
di: Chaar, Mohamad Mofeed, et al.
Pubblicazione: (2025)
Surgical Depth Anything: Depth Estimation for Surgical Scenes using Foundation Models
di: Lou, Ange, et al.
Pubblicazione: (2024)
di: Lou, Ange, et al.
Pubblicazione: (2024)
Feature Mixing Approach for Detecting Intraoperative Adverse Events in Laparoscopic Roux-en-Y Gastric Bypass Surgery
di: Bose, Rupak, et al.
Pubblicazione: (2025)
di: Bose, Rupak, et al.
Pubblicazione: (2025)
Boosting Omnidirectional Stereo Matching with a Pre-trained Depth Foundation Model
di: Endres, Jannik, et al.
Pubblicazione: (2025)
di: Endres, Jannik, et al.
Pubblicazione: (2025)
Depth Attention for Robust RGB Tracking
di: Liu, Yu, et al.
Pubblicazione: (2024)
di: Liu, Yu, et al.
Pubblicazione: (2024)
Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?
di: Che, Chengan, et al.
Pubblicazione: (2026)
di: Che, Chengan, et al.
Pubblicazione: (2026)
VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones
di: Shen, Lefei, et al.
Pubblicazione: (2025)
di: Shen, Lefei, et al.
Pubblicazione: (2025)
RGB-Event HyperGraph Prompt for Kilometer Marker Recognition based on Pre-trained Foundation Models
di: Xian, Xiaoyu, et al.
Pubblicazione: (2026)
di: Xian, Xiaoyu, et al.
Pubblicazione: (2026)
Self-supervised Learning via Cluster Distance Prediction for Operating Room Context Awareness
di: Hamoud, Idris, et al.
Pubblicazione: (2024)
di: Hamoud, Idris, et al.
Pubblicazione: (2024)
3D Scene Graph Guided Vision-Language Pre-training
di: Liu, Hao, et al.
Pubblicazione: (2024)
di: Liu, Hao, et al.
Pubblicazione: (2024)
NimbleD: Enhancing Self-supervised Monocular Depth Estimation with Pseudo-labels and Large-scale Video Pre-training
di: Luginov, Albert, et al.
Pubblicazione: (2024)
di: Luginov, Albert, et al.
Pubblicazione: (2024)
Computer Vision Foundation Models in Endoscopy: Proof of Concept in Oropharyngeal Cancer
di: Alberto Paderno, et al.
Pubblicazione: (2024)
di: Alberto Paderno, et al.
Pubblicazione: (2024)
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
di: Chen, Yitong, et al.
Pubblicazione: (2025)
di: Chen, Yitong, et al.
Pubblicazione: (2025)
Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence
di: Rau, Anita, et al.
Pubblicazione: (2025)
di: Rau, Anita, et al.
Pubblicazione: (2025)
Depth-guided NeRF Training via Earth Mover's Distance
di: Rau, Anita, et al.
Pubblicazione: (2024)
di: Rau, Anita, et al.
Pubblicazione: (2024)
Fully Exploiting Vision Foundation Model's Profound Prior Knowledge for Generalizable RGB-Depth Driving Scene Parsing
di: Guo, Sicen, et al.
Pubblicazione: (2025)
di: Guo, Sicen, et al.
Pubblicazione: (2025)
An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model
di: Tian, Yuxin, et al.
Pubblicazione: (2024)
di: Tian, Yuxin, et al.
Pubblicazione: (2024)
A Positive Answer to a Question of K. Borsuk on the Capacity of Polyhedra with Finite by Cyclic Fundamental Group
di: Mohareri, Mojtaba, et al.
Pubblicazione: (2023)
di: Mohareri, Mojtaba, et al.
Pubblicazione: (2023)
EndoPBR: Material and Lighting Estimation for Photorealistic Surgical Simulations via Physically-based Rendering
di: Han, John J., et al.
Pubblicazione: (2025)
di: Han, John J., et al.
Pubblicazione: (2025)
Documenti analoghi
-
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
di: Perez, Alejandra, et al.
Pubblicazione: (2026) -
Rethinking RGB-D Fusion for Semantic Segmentation in Surgical Datasets
di: Jamal, Muhammad Abdullah, et al.
Pubblicazione: (2024) -
SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
di: Perez, Alejandra, et al.
Pubblicazione: (2025) -
A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders
di: Jamal, Muhammad Abdullah, et al.
Pubblicazione: (2024) -
Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
di: Venkatesh, Danush Kumar, et al.
Pubblicazione: (2025)