HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Biagini, Diego, Navab, Nassir, Farshad, Azade |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary
by: He, Jingyi, et al.
Published: (2026)
by: He, Jingyi, et al.
Published: (2026)
SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction
by: Köksal, Çağhan, et al.
Published: (2024)
by: Köksal, Çağhan, et al.
Published: (2024)
Conformable Convolution for Topologically Aware Learning of Complex Anatomical Structures
by: Yeganeh, Yousef, et al.
Published: (2024)
by: Yeganeh, Yousef, et al.
Published: (2024)
Physics-Informed Latent Diffusion for Multimodal Brain MRI Synthesis
by: Lüpke, Sven, et al.
Published: (2024)
by: Lüpke, Sven, et al.
Published: (2024)
Stress-Aware Resilient Neural Training
by: Shakarami, Ashkan, et al.
Published: (2025)
by: Shakarami, Ashkan, et al.
Published: (2025)
Millimeter-wave Imaging for Anthropometric Body Measurement
by: Senne, Miriam, et al.
Published: (2026)
by: Senne, Miriam, et al.
Published: (2026)
SURGIVID: Annotation-Efficient Surgical Video Object Discovery
by: Köksal, Çağhan, et al.
Published: (2024)
by: Köksal, Çağhan, et al.
Published: (2024)
Latent Drifting in Diffusion Models for Counterfactual Medical Image Synthesis
by: Yeganeh, Yousef, et al.
Published: (2024)
by: Yeganeh, Yousef, et al.
Published: (2024)
DeepAf: One-Shot Spatiospectral Auto-Focus Model for Digital Pathology
by: Yeganeh, Yousef, et al.
Published: (2025)
by: Yeganeh, Yousef, et al.
Published: (2025)
VISAGE: Video Synthesis using Action Graphs for Surgery
by: Yeganeh, Yousef, et al.
Published: (2024)
by: Yeganeh, Yousef, et al.
Published: (2024)
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
by: Yuan, Kun, et al.
Published: (2024)
by: Yuan, Kun, et al.
Published: (2024)
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
VeLU: Variance-enhanced Learning Unit for Deep Neural Networks
by: Shakarami, Ashkan, et al.
Published: (2025)
by: Shakarami, Ashkan, et al.
Published: (2025)
SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos
by: Wu, Jinlin, et al.
Published: (2026)
by: Wu, Jinlin, et al.
Published: (2026)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
by: Wang, Guankun, et al.
Published: (2025)
by: Wang, Guankun, et al.
Published: (2025)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
by: Yuan, Kun, et al.
Published: (2024)
by: Yuan, Kun, et al.
Published: (2024)
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
by: Stilz, Florian, et al.
Published: (2026)
by: Stilz, Florian, et al.
Published: (2026)
Mitigating Biases in Surgical Operating Rooms with Geometry
by: Wang, Tony Danjun, et al.
Published: (2025)
by: Wang, Tony Danjun, et al.
Published: (2025)
Next-generation Surgical Navigation: Marker-less Multi-view 6DoF Pose Estimation of Surgical Instruments
by: Hein, Jonas, et al.
Published: (2023)
by: Hein, Jonas, et al.
Published: (2023)
Unit-Based Histopathology Tissue Segmentation via Multi-Level Feature Representation
by: Shakarami, Ashkan, et al.
Published: (2025)
by: Shakarami, Ashkan, et al.
Published: (2025)
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
by: Huang, Yiming, et al.
Published: (2025)
by: Huang, Yiming, et al.
Published: (2025)
From Linear Probing to Joint-Weighted Token Hierarchy: A Foundation Model Bridging Global and Cellular Representations in Biomarker Detection
by: Liu, Jingsong, et al.
Published: (2025)
by: Liu, Jingsong, et al.
Published: (2025)
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
by: Holm, Felix, et al.
Published: (2025)
by: Holm, Felix, et al.
Published: (2025)
Hybrid Functional Maps for Crease-Aware Non-Isometric Shape Matching
by: Bastian, Lennart, et al.
Published: (2023)
by: Bastian, Lennart, et al.
Published: (2023)
Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis
by: Chen, Tingxuan, et al.
Published: (2025)
by: Chen, Tingxuan, et al.
Published: (2025)
Diffusion as Sound Propagation: Physics-inspired Model for Ultrasound Image Generation
by: Domínguez, Marina, et al.
Published: (2024)
by: Domínguez, Marina, et al.
Published: (2024)
Temporal Differential Fields for 4D Motion Modeling via Image-to-Video Synthesis
by: You, Xin, et al.
Published: (2025)
by: You, Xin, et al.
Published: (2025)
HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models
by: Guo, Yansong, et al.
Published: (2026)
by: Guo, Yansong, et al.
Published: (2026)
Beyond Role-Based Surgical Domain Modeling: Generalizable Re-Identification in the Operating Room
by: Wang, Tony Danjun, et al.
Published: (2025)
by: Wang, Tony Danjun, et al.
Published: (2025)
How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
by: An, Joungbin, et al.
Published: (2025)
by: An, Joungbin, et al.
Published: (2025)
ORacle: Large Vision-Language Models for Knowledge-Guided Holistic OR Domain Modeling
by: Özsoy, Ege, et al.
Published: (2024)
by: Özsoy, Ege, et al.
Published: (2024)
BridgeSplat: Bidirectionally Coupled CT and Non-Rigid Gaussian Splatting for Deformable Intraoperative Surgical Navigation
by: Fehrentz, Maximilian, et al.
Published: (2025)
by: Fehrentz, Maximilian, et al.
Published: (2025)
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
by: Guo, Diandian, et al.
Published: (2026)
by: Guo, Diandian, et al.
Published: (2026)
SurgPETL: Parameter-Efficient Image-to-Surgical-Video Transfer Learning for Surgical Phase Recognition
by: Yang, Shu, et al.
Published: (2024)
by: Yang, Shu, et al.
Published: (2024)
SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy
by: Li, Shi, et al.
Published: (2026)
by: Li, Shi, et al.
Published: (2026)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
by: Yuan, Kun, et al.
Published: (2023)
by: Yuan, Kun, et al.
Published: (2023)
HieraFashDiff: Hierarchical Fashion Design with Multi-stage Diffusion Models
by: Xie, Zhifeng, et al.
Published: (2024)
by: Xie, Zhifeng, et al.
Published: (2024)
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
by: Özsoy, Ege, et al.
Published: (2025)
by: Özsoy, Ege, et al.
Published: (2025)
Similar Items
-
SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary
by: He, Jingyi, et al.
Published: (2026) -
SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction
by: Köksal, Çağhan, et al.
Published: (2024) -
Conformable Convolution for Topologically Aware Learning of Complex Anatomical Structures
by: Yeganeh, Yousef, et al.
Published: (2024) -
Physics-Informed Latent Diffusion for Multimodal Brain MRI Synthesis
by: Lüpke, Sven, et al.
Published: (2024) -
Stress-Aware Resilient Neural Training
by: Shakarami, Ashkan, et al.
Published: (2025)