SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Jingyi, Zhou, Yue, Bai, Long, Yuan, Kun, Navab, Nassir, Bi, Yuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation
von: Biagini, Diego, et al.
Veröffentlicht: (2025)
von: Biagini, Diego, et al.
Veröffentlicht: (2025)
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
SURGIVID: Annotation-Efficient Surgical Video Object Discovery
von: Köksal, Çağhan, et al.
Veröffentlicht: (2024)
von: Köksal, Çağhan, et al.
Veröffentlicht: (2024)
Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis
von: Chen, Tingxuan, et al.
Veröffentlicht: (2025)
von: Chen, Tingxuan, et al.
Veröffentlicht: (2025)
UltraAD: Fine-Grained Ultrasound Anomaly Classification via Few-Shot CLIP Adaptation
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos
von: Wu, Jinlin, et al.
Veröffentlicht: (2026)
von: Wu, Jinlin, et al.
Veröffentlicht: (2026)
Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion
von: Wei, Meng, et al.
Veröffentlicht: (2026)
von: Wei, Meng, et al.
Veröffentlicht: (2026)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
von: Chen, Tong, et al.
Veröffentlicht: (2024)
von: Chen, Tong, et al.
Veröffentlicht: (2024)
Advancing Surgical VQA with Scene Graph Knowledge
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
von: Yuan, Kun, et al.
Veröffentlicht: (2023)
SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction
von: Köksal, Çağhan, et al.
Veröffentlicht: (2024)
von: Köksal, Çağhan, et al.
Veröffentlicht: (2024)
SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy
von: Li, Shi, et al.
Veröffentlicht: (2026)
von: Li, Shi, et al.
Veröffentlicht: (2026)
Class-Aware Cartilage Segmentation for Autonomous US-CT Registration in Robotic Intercostal Ultrasound Imaging
von: Jiang, Zhongliang, et al.
Veröffentlicht: (2024)
von: Jiang, Zhongliang, et al.
Veröffentlicht: (2024)
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
von: Zhou, Yue, et al.
Veröffentlicht: (2026)
von: Zhou, Yue, et al.
Veröffentlicht: (2026)
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
von: Stilz, Florian, et al.
Veröffentlicht: (2026)
von: Stilz, Florian, et al.
Veröffentlicht: (2026)
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
von: Rohrmoser, Nikolo, et al.
Veröffentlicht: (2026)
von: Rohrmoser, Nikolo, et al.
Veröffentlicht: (2026)
Mitigating Biases in Surgical Operating Rooms with Geometry
von: Wang, Tony Danjun, et al.
Veröffentlicht: (2025)
von: Wang, Tony Danjun, et al.
Veröffentlicht: (2025)
Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
Robotic Ultrasound Makes CBCT Alive
von: Li, Feng, et al.
Veröffentlicht: (2026)
von: Li, Feng, et al.
Veröffentlicht: (2026)
Hybrid Functional Maps for Crease-Aware Non-Isometric Shape Matching
von: Bastian, Lennart, et al.
Veröffentlicht: (2023)
von: Bastian, Lennart, et al.
Veröffentlicht: (2023)
Specialized Foundation Models for Intelligent Operating Rooms
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
From Linear Probing to Joint-Weighted Token Hierarchy: A Foundation Model Bridging Global and Cellular Representations in Biomarker Detection
von: Liu, Jingsong, et al.
Veröffentlicht: (2025)
von: Liu, Jingsong, et al.
Veröffentlicht: (2025)
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
von: Hu, Ming, et al.
Veröffentlicht: (2024)
von: Hu, Ming, et al.
Veröffentlicht: (2024)
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
von: Holm, Felix, et al.
Veröffentlicht: (2025)
von: Holm, Felix, et al.
Veröffentlicht: (2025)
LapFM: A Laparoscopic Segmentation Foundation Model via Hierarchical Concept Evolving Pre-training
von: Xu, Qing, et al.
Veröffentlicht: (2025)
von: Xu, Qing, et al.
Veröffentlicht: (2025)
MI-SegNet: Mutual Information-Based US Segmentation for Unseen Domain Generalization
von: Bi, Yuan, et al.
Veröffentlicht: (2023)
von: Bi, Yuan, et al.
Veröffentlicht: (2023)
BridgeSplat: Bidirectionally Coupled CT and Non-Rigid Gaussian Splatting for Deformable Intraoperative Surgical Navigation
von: Fehrentz, Maximilian, et al.
Veröffentlicht: (2025)
von: Fehrentz, Maximilian, et al.
Veröffentlicht: (2025)
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
von: Guo, Diandian, et al.
Veröffentlicht: (2026)
von: Guo, Diandian, et al.
Veröffentlicht: (2026)
SurgPETL: Parameter-Efficient Image-to-Surgical-Video Transfer Learning for Surgical Phase Recognition
von: Yang, Shu, et al.
Veröffentlicht: (2024)
von: Yang, Shu, et al.
Veröffentlicht: (2024)
Bridging Vision and Language for Robust Context-Aware Surgical Point Tracking: The VL-SurgPT Dataset and Benchmark
von: Zhou, Rulin, et al.
Veröffentlicht: (2025)
von: Zhou, Rulin, et al.
Veröffentlicht: (2025)
Temporal Differential Fields for 4D Motion Modeling via Image-to-Video Synthesis
von: You, Xin, et al.
Veröffentlicht: (2025)
von: You, Xin, et al.
Veröffentlicht: (2025)
SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark
von: Wang, Gui, et al.
Veröffentlicht: (2026)
von: Wang, Gui, et al.
Veröffentlicht: (2026)
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
SemiVT-Surge: Semi-Supervised Video Transformer for Surgical Phase Recognition
von: Li, Yiping, et al.
Veröffentlicht: (2025)
von: Li, Yiping, et al.
Veröffentlicht: (2025)
Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025
von: Zia, Aneeq, et al.
Veröffentlicht: (2023)
von: Zia, Aneeq, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation
von: Biagini, Diego, et al.
Veröffentlicht: (2025) -
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
von: Yuan, Kun, et al.
Veröffentlicht: (2024) -
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
von: Yuan, Kun, et al.
Veröffentlicht: (2024) -
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
von: Wang, Guankun, et al.
Veröffentlicht: (2025) -
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
von: Chen, Zhen, et al.
Veröffentlicht: (2025)