Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Tingxuan, Yuan, Kun, Srivastav, Vinkle, Navab, Nassir, Padoy, Nicolas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
von: Stilz, Florian, et al.
Veröffentlicht: (2026)
von: Stilz, Florian, et al.
Veröffentlicht: (2026)
Advancing Surgical VQA with Scene Graph Knowledge
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
von: Walimbe, Soham, et al.
Veröffentlicht: (2025)
von: Walimbe, Soham, et al.
Veröffentlicht: (2025)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
Multi-modal Representations for Fine-grained Multi-label Critical View of Safety Recognition
von: Baby, Britty, et al.
Veröffentlicht: (2025)
von: Baby, Britty, et al.
Veröffentlicht: (2025)
SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction
von: Köksal, Çağhan, et al.
Veröffentlicht: (2024)
von: Köksal, Çağhan, et al.
Veröffentlicht: (2024)
SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose Estimation
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2024)
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2024)
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
von: Holm, Felix, et al.
Veröffentlicht: (2025)
von: Holm, Felix, et al.
Veröffentlicht: (2025)
Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion
von: Wei, Meng, et al.
Veröffentlicht: (2026)
von: Wei, Meng, et al.
Veröffentlicht: (2026)
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy
von: Li, Shi, et al.
Veröffentlicht: (2026)
von: Li, Shi, et al.
Veröffentlicht: (2026)
End-to-End Learning of Multi-Organ Implicit Surfaces from 3D Medical Imaging Data
von: Zarin, Farahdiba, et al.
Veröffentlicht: (2025)
von: Zarin, Farahdiba, et al.
Veröffentlicht: (2025)
Overcoming Dimensional Collapse in Self-supervised Contrastive Learning for Medical Image Segmentation
von: Hassanpour, Jamshid, et al.
Veröffentlicht: (2024)
von: Hassanpour, Jamshid, et al.
Veröffentlicht: (2024)
Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
von: Zhai, Guangyao, et al.
Veröffentlicht: (2025)
von: Zhai, Guangyao, et al.
Veröffentlicht: (2025)
How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition
von: Hamoud, Idris, et al.
Veröffentlicht: (2025)
von: Hamoud, Idris, et al.
Veröffentlicht: (2025)
From Linear Probing to Joint-Weighted Token Hierarchy: A Foundation Model Bridging Global and Cellular Representations in Biomarker Detection
von: Liu, Jingsong, et al.
Veröffentlicht: (2025)
von: Liu, Jingsong, et al.
Veröffentlicht: (2025)
Endoshare: A Publicly Available, Surgeons-Friendly Solution to De-Identify and Manage Surgical Videos
von: Arboit, Lorenzo, et al.
Veröffentlicht: (2025)
von: Arboit, Lorenzo, et al.
Veröffentlicht: (2025)
Jumpstarting Surgical Computer Vision
von: Alapatt, Deepak, et al.
Veröffentlicht: (2023)
von: Alapatt, Deepak, et al.
Veröffentlicht: (2023)
Few-shot Writer Adaptation via Multimodal In-Context Learning
von: Simon, Tom, et al.
Veröffentlicht: (2026)
von: Simon, Tom, et al.
Veröffentlicht: (2026)
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
Self-Supervised Uncalibrated Multi-View Video Anonymization in the Operating Room
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
von: Chen, Keqi, et al.
Veröffentlicht: (2026)
DeepAf: One-Shot Spatiospectral Auto-Focus Model for Digital Pathology
von: Yeganeh, Yousef, et al.
Veröffentlicht: (2025)
von: Yeganeh, Yousef, et al.
Veröffentlicht: (2025)
Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation
von: Bahram, Yara, et al.
Veröffentlicht: (2025)
von: Bahram, Yara, et al.
Veröffentlicht: (2025)
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
Siamese Transformer Networks for Few-shot Image Classification
von: Jiang, Weihao, et al.
Veröffentlicht: (2024)
von: Jiang, Weihao, et al.
Veröffentlicht: (2024)
FAD: Frequency Adaptation and Diversion for Cross-domain Few-shot Learning
von: Shi, Ruixiao, et al.
Veröffentlicht: (2025)
von: Shi, Ruixiao, et al.
Veröffentlicht: (2025)
Robotic Ultrasound Makes CBCT Alive
von: Li, Feng, et al.
Veröffentlicht: (2026)
von: Li, Feng, et al.
Veröffentlicht: (2026)
CholecTrack20: A Multi-Perspective Tracking Dataset for Surgical Tools
von: Nwoye, Chinedu Innocent, et al.
Veröffentlicht: (2023)
von: Nwoye, Chinedu Innocent, et al.
Veröffentlicht: (2023)
A Skull-Adaptive Framework for AI-Based 3D Transcranial Focused Ultrasound Simulation
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2025)
von: Srivastav, Vinkle, et al.
Veröffentlicht: (2025)
Speckle2Self: Self-Supervised Ultrasound Speckle Reduction Without Clean Data
von: Li, Xuesong, et al.
Veröffentlicht: (2025)
von: Li, Xuesong, et al.
Veröffentlicht: (2025)
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
von: Hu, Ming, et al.
Veröffentlicht: (2024)
von: Hu, Ming, et al.
Veröffentlicht: (2024)
Artificial Intelligence for the Assessment of Peritoneal Carcinosis during Diagnostic Laparoscopy for Advanced Ovarian Cancer
von: Oliva, Riccardo, et al.
Veröffentlicht: (2025)
von: Oliva, Riccardo, et al.
Veröffentlicht: (2025)
CAT-SG: A Large Dynamic Scene Graph Dataset for Fine-Grained Understanding of Cataract Surgery
von: Holm, Felix, et al.
Veröffentlicht: (2025)
von: Holm, Felix, et al.
Veröffentlicht: (2025)
Energy-based Tissue Manifolds for Longitudinal Multiparametric MRI Analysis
von: Tehlan, Kartikay, et al.
Veröffentlicht: (2026)
von: Tehlan, Kartikay, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024) -
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
von: Yuan, Kun, et al.
Veröffentlicht: (2024) -
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
von: Stilz, Florian, et al.
Veröffentlicht: (2026) -
Advancing Surgical VQA with Scene Graph Knowledge
von: Yuan, Kun, et al.
Veröffentlicht: (2023) -
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
von: Walimbe, Soham, et al.
Veröffentlicht: (2025)