Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Kun, Chen, Tingxuan, Li, Shi, Lavanchy, Joel L., Heiliger, Christian, Özsoy, Ege, Huang, Yiming, Bai, Long, Navab, Nassir, Srivastav, Vinkle, Ren, Hongliang, Padoy, Nicolas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis
von: Chen, Tingxuan, et al.
Veröffentlicht: (2025)
von: Chen, Tingxuan, et al.
Veröffentlicht: (2025)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
Advancing Surgical VQA with Scene Graph Knowledge
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
von: Stilz, Florian, et al.
Veröffentlicht: (2026)
von: Stilz, Florian, et al.
Veröffentlicht: (2026)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
von: Walimbe, Soham, et al.
Veröffentlicht: (2025)
von: Walimbe, Soham, et al.
Veröffentlicht: (2025)
BridgeSplat: Bidirectionally Coupled CT and Non-Rigid Gaussian Splatting for Deformable Intraoperative Surgical Navigation
von: Fehrentz, Maximilian, et al.
Veröffentlicht: (2025)
von: Fehrentz, Maximilian, et al.
Veröffentlicht: (2025)
ORacle: Large Vision-Language Models for Knowledge-Guided Holistic OR Domain Modeling
von: Özsoy, Ege, et al.
Veröffentlicht: (2024)
von: Özsoy, Ege, et al.
Veröffentlicht: (2024)
PhenoKG: Knowledge Graph-Driven Gene Discovery and Patient Insights from Phenotypes Alone
von: Zaripova, Kamilia, et al.
Veröffentlicht: (2025)
von: Zaripova, Kamilia, et al.
Veröffentlicht: (2025)
SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose Estimation
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2024)
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2024)
PanORama: Multiview Consistent Panoptic Segmentation in Operating Rooms
von: Gürbüz, Tuna, et al.
Veröffentlicht: (2026)
von: Gürbüz, Tuna, et al.
Veröffentlicht: (2026)
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
When do they StOP?: A First Step Towards Automatically Identifying Team Communication in the Operating Room
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
Prototype-Based Knowledge Guidance for Fine-Grained Structured Radiology Reporting
von: Pellegrini, Chantal, et al.
Veröffentlicht: (2026)
von: Pellegrini, Chantal, et al.
Veröffentlicht: (2026)
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
von: Pellegrini, Chantal, et al.
Veröffentlicht: (2023)
von: Pellegrini, Chantal, et al.
Veröffentlicht: (2023)
Beyond Role-Based Surgical Domain Modeling: Generalizable Re-Identification in the Operating Room
von: Wang, Tony Danjun, et al.
Veröffentlicht: (2025)
von: Wang, Tony Danjun, et al.
Veröffentlicht: (2025)
Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
End-to-End Learning of Multi-Organ Implicit Surfaces from 3D Medical Imaging Data
von: Zarin, Farahdiba, et al.
Veröffentlicht: (2025)
von: Zarin, Farahdiba, et al.
Veröffentlicht: (2025)
Overcoming Dimensional Collapse in Self-supervised Contrastive Learning for Medical Image Segmentation
von: Hassanpour, Jamshid, et al.
Veröffentlicht: (2024)
von: Hassanpour, Jamshid, et al.
Veröffentlicht: (2024)
Specialized Foundation Models for Intelligent Operating Rooms
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning
von: Bani-Harouni, David, et al.
Veröffentlicht: (2025)
von: Bani-Harouni, David, et al.
Veröffentlicht: (2025)
EHR2Path: Scalable Modeling of Longitudinal Patient Pathways from Multimodal Electronic Health Records
von: Pellegrini, Chantal, et al.
Veröffentlicht: (2025)
von: Pellegrini, Chantal, et al.
Veröffentlicht: (2025)
Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition
von: Hamoud, Idris, et al.
Veröffentlicht: (2025)
von: Hamoud, Idris, et al.
Veröffentlicht: (2025)
Jumpstarting Surgical Computer Vision
von: Alapatt, Deepak, et al.
Veröffentlicht: (2023)
von: Alapatt, Deepak, et al.
Veröffentlicht: (2023)
Endoshare: A Publicly Available, Surgeons-Friendly Solution to De-Identify and Manage Surgical Videos
von: Arboit, Lorenzo, et al.
Veröffentlicht: (2025)
von: Arboit, Lorenzo, et al.
Veröffentlicht: (2025)
Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion
von: Wei, Meng, et al.
Veröffentlicht: (2026)
von: Wei, Meng, et al.
Veröffentlicht: (2026)
TrackOR: Towards Personalized Intelligent Operating Rooms Through Robust Tracking
von: Wang, Tony Danjun, et al.
Veröffentlicht: (2025)
von: Wang, Tony Danjun, et al.
Veröffentlicht: (2025)
MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
Multi-modal Representations for Fine-grained Multi-label Critical View of Safety Recognition
von: Baby, Britty, et al.
Veröffentlicht: (2025)
von: Baby, Britty, et al.
Veröffentlicht: (2025)
SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy
von: Li, Shi, et al.
Veröffentlicht: (2026)
von: Li, Shi, et al.
Veröffentlicht: (2026)
UltraAD: Fine-Grained Ultrasound Anomaly Classification via Few-Shot CLIP Adaptation
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
Location-Free Scene Graph Generation
von: Özsoy, Ege, et al.
Veröffentlicht: (2023)
von: Özsoy, Ege, et al.
Veröffentlicht: (2023)
Self-Supervised Uncalibrated Multi-View Video Anonymization in the Operating Room
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models
von: Bani-Harouni, David, et al.
Veröffentlicht: (2025)
von: Bani-Harouni, David, et al.
Veröffentlicht: (2025)
CholecTrack20: A Multi-Perspective Tracking Dataset for Surgical Tools
von: Nwoye, Chinedu Innocent, et al.
Veröffentlicht: (2023)
von: Nwoye, Chinedu Innocent, et al.
Veröffentlicht: (2023)
SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary
von: He, Jingyi, et al.
Veröffentlicht: (2026)
von: He, Jingyi, et al.
Veröffentlicht: (2026)
A Skull-Adaptive Framework for AI-Based 3D Transcranial Focused Ultrasound Simulation
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2025)
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis
von: Chen, Tingxuan, et al.
Veröffentlicht: (2025) -
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024) -
Advancing Surgical VQA with Scene Graph Knowledge
von: Yuan, Kun, et al.
Veröffentlicht: (2023) -
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
von: Yuan, Kun, et al.
Veröffentlicht: (2024) -
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
von: Stilz, Florian, et al.
Veröffentlicht: (2026)