Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hamoud, Idris, Srivastav, Vinkle, Jamal, Muhammad Abdullah, Mutter, Didier, Mohareri, Omid, Padoy, Nicolas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Supervised Uncalibrated Multi-View Video Anonymization in the Operating Room
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
Self-supervised Learning via Cluster Distance Prediction for Operating Room Context Awareness
von: Hamoud, Idris, et al.
Veröffentlicht: (2024)
von: Hamoud, Idris, et al.
Veröffentlicht: (2024)
Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
Overcoming Dimensional Collapse in Self-supervised Contrastive Learning for Medical Image Segmentation
von: Hassanpour, Jamshid, et al.
Veröffentlicht: (2024)
von: Hassanpour, Jamshid, et al.
Veröffentlicht: (2024)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
von: Stilz, Florian, et al.
Veröffentlicht: (2026)
von: Stilz, Florian, et al.
Veröffentlicht: (2026)
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose Estimation
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2024)
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2024)
Rethinking RGB-D Fusion for Semantic Segmentation in Surgical Datasets
von: Jamal, Muhammad Abdullah, et al.
Veröffentlicht: (2024)
von: Jamal, Muhammad Abdullah, et al.
Veröffentlicht: (2024)
When do they StOP?: A First Step Towards Automatically Identifying Team Communication in the Operating Room
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders
von: Jamal, Muhammad Abdullah, et al.
Veröffentlicht: (2024)
von: Jamal, Muhammad Abdullah, et al.
Veröffentlicht: (2024)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
von: Walimbe, Soham, et al.
Veröffentlicht: (2025)
von: Walimbe, Soham, et al.
Veröffentlicht: (2025)
End-to-End Learning of Multi-Organ Implicit Surfaces from 3D Medical Imaging Data
von: Zarin, Farahdiba, et al.
Veröffentlicht: (2025)
von: Zarin, Farahdiba, et al.
Veröffentlicht: (2025)
Where are they looking in the operating room?
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis
von: Chen, Tingxuan, et al.
Veröffentlicht: (2025)
von: Chen, Tingxuan, et al.
Veröffentlicht: (2025)
Endoshare: A Publicly Available, Surgeons-Friendly Solution to De-Identify and Manage Surgical Videos
von: Arboit, Lorenzo, et al.
Veröffentlicht: (2025)
von: Arboit, Lorenzo, et al.
Veröffentlicht: (2025)
VidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery
von: Honarmand, Mohammadmahdi, et al.
Veröffentlicht: (2024)
von: Honarmand, Mohammadmahdi, et al.
Veröffentlicht: (2024)
DExTeR: Weakly Semi-Supervised Object Detection with Class and Instance Experts for Medical Imaging
von: Meyer, Adrien, et al.
Veröffentlicht: (2026)
von: Meyer, Adrien, et al.
Veröffentlicht: (2026)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
Multi-modal Representations for Fine-grained Multi-label Critical View of Safety Recognition
von: Baby, Britty, et al.
Veröffentlicht: (2025)
von: Baby, Britty, et al.
Veröffentlicht: (2025)
SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
von: Perez, Alejandra, et al.
Veröffentlicht: (2025)
von: Perez, Alejandra, et al.
Veröffentlicht: (2025)
State-Change Learning for Prediction of Future Events in Endoscopic Videos
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
Advancing Surgical VQA with Scene Graph Knowledge
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
Jumpstarting Surgical Computer Vision
von: Alapatt, Deepak, et al.
Veröffentlicht: (2023)
von: Alapatt, Deepak, et al.
Veröffentlicht: (2023)
SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy
von: Li, Shi, et al.
Veröffentlicht: (2026)
von: Li, Shi, et al.
Veröffentlicht: (2026)
CycleSAM: Few-Shot Surgical Scene Segmentation with Cycle- and Scene-Consistent Feature Matching
von: Murali, Aditya, et al.
Veröffentlicht: (2024)
von: Murali, Aditya, et al.
Veröffentlicht: (2024)
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
von: Perez, Alejandra, et al.
Veröffentlicht: (2026)
von: Perez, Alejandra, et al.
Veröffentlicht: (2026)
AdaEmbed: Semi-supervised Domain Adaptation in the Embedding Space
von: Mottaghi, Ali, et al.
Veröffentlicht: (2024)
von: Mottaghi, Ali, et al.
Veröffentlicht: (2024)
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
von: Han, John J., et al.
Veröffentlicht: (2026)
von: Han, John J., et al.
Veröffentlicht: (2026)
SurgiTrack: Fine-Grained Multi-Class Multi-Tool Tracking in Surgical Videos
von: Nwoye, Chinedu Innocent, et al.
Veröffentlicht: (2024)
von: Nwoye, Chinedu Innocent, et al.
Veröffentlicht: (2024)
A Skull-Adaptive Framework for AI-Based 3D Transcranial Focused Ultrasound Simulation
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2025)
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2025)
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
von: Hu, Ming, et al.
Veröffentlicht: (2024)
von: Hu, Ming, et al.
Veröffentlicht: (2024)
UltraSam: A Foundation Model for Ultrasound using Large Open-Access Segmentation Datasets
von: Meyer, Adrien, et al.
Veröffentlicht: (2024)
von: Meyer, Adrien, et al.
Veröffentlicht: (2024)
Early Operative Difficulty Assessment in Laparoscopic Cholecystectomy via Snapshot-Centric Video Analysis
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
S4M: 4-points to Segment Anything
von: Meyer, Adrien, et al.
Veröffentlicht: (2025)
von: Meyer, Adrien, et al.
Veröffentlicht: (2025)
Seeing Through Smoke: Surgical Desmoking for Improved Visual Perception
von: Lu, Jingpei, et al.
Veröffentlicht: (2026)
von: Lu, Jingpei, et al.
Veröffentlicht: (2026)
Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
von: Venkatesh, Danush Kumar, et al.
Veröffentlicht: (2025)
von: Venkatesh, Danush Kumar, et al.
Veröffentlicht: (2025)
Surgical Tattoos in Infrared: A Dataset for Quantifying Tissue Tracking and Mapping
von: Schmidt, Adam, et al.
Veröffentlicht: (2023)
von: Schmidt, Adam, et al.
Veröffentlicht: (2023)
Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Self-Supervised Uncalibrated Multi-View Video Anonymization in the Operating Room
von: Chen, Keqi, et al.
Veröffentlicht: (2026) -
Self-supervised Learning via Cluster Distance Prediction for Operating Room Context Awareness
von: Hamoud, Idris, et al.
Veröffentlicht: (2024) -
Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes
von: Chen, Keqi, et al.
Veröffentlicht: (2025) -
Overcoming Dimensional Collapse in Self-supervised Contrastive Learning for Medical Image Segmentation
von: Hassanpour, Jamshid, et al.
Veröffentlicht: (2024) -
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024)