SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Rapuri, Sampath, Seenivasan, Lalithkumar, Schneider, Dominik, Soberanis-Mukul, Roger, He, Yufan, Ding, Hao, Xu, Jiru, Yu, Chenhao, Jing, Chenyan, Guo, Pengfei, Xu, Daguang, Unberath, Mathias |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
by: Maksutova, Aiza, et al.
Published: (2026)
by: Maksutova, Aiza, et al.
Published: (2026)
Beyond Rigid AI: Towards Natural Human-Machine Symbiosis for Interoperative Surgical Assistance
by: Seenivasan, Lalithkumar, et al.
Published: (2025)
by: Seenivasan, Lalithkumar, et al.
Published: (2025)
Towards Controllable Video Synthesis of Routine and Rare OR Events
by: Schneider, Dominik, et al.
Published: (2026)
by: Schneider, Dominik, et al.
Published: (2026)
From Generalization to Precision: Exploring SAM for Tool Segmentation in Surgical Environments
by: Oguine, Kanyifeechukwu J., et al.
Published: (2024)
by: Oguine, Kanyifeechukwu J., et al.
Published: (2024)
Online Reasoning Video Segmentation with Just-in-Time Digital Twins
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
Position: Foundation Models Need Digital Twin Representations
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
BronchOpt : Vision-Based Pose Optimization with Fine-Tuned Foundation Models for Accurate Bronchoscopy Navigation
by: Shu, Hongchao, et al.
Published: (2025)
by: Shu, Hongchao, et al.
Published: (2025)
Investigating a Policy-Based Formulation for Endoscopic Camera Pose Recovery
by: Mangulabnan, Jan Emily, et al.
Published: (2026)
by: Mangulabnan, Jan Emily, et al.
Published: (2026)
Privacy-Preserving Operating Room Workflow Analysis using Digital Twins
by: Perez, Alejandra, et al.
Published: (2025)
by: Perez, Alejandra, et al.
Published: (2025)
Did you just see that? Arbitrary view synthesis for egocentric replay of operating room workflows from ambient sensors
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
Constrained Natural Language Action Planning for Resilient Embodied Systems
by: Byrd, Grayson, et al.
Published: (2025)
by: Byrd, Grayson, et al.
Published: (2025)
MoSFormer: Augmenting Temporal Context with Memory of Surgery for Surgical Phase Recognition
by: Ding, Hao, et al.
Published: (2025)
by: Ding, Hao, et al.
Published: (2025)
Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models
by: Ding, Hao, et al.
Published: (2024)
by: Ding, Hao, et al.
Published: (2024)
TwinOR: Photorealistic Digital Twins of Dynamic Operating Rooms for Embodied AI Research
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
Investigating Robot Control Policy Learning for Autonomous X-ray-guided Spine Procedures
by: Klitzner, Florence, et al.
Published: (2025)
by: Klitzner, Florence, et al.
Published: (2025)
Performance and Non-adversarial Robustness of the Segment Anything Model 2 in Surgical Video Segmentation
by: Shen, Yiqing, et al.
Published: (2024)
by: Shen, Yiqing, et al.
Published: (2024)
Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling
by: He, Yufan, et al.
Published: (2025)
by: He, Yufan, et al.
Published: (2025)
Seamless Augmented Reality Integration in Arthroscopy: A Pipeline for Articular Reconstruction and Guidance
by: Shu, Hongchao, et al.
Published: (2024)
by: Shu, Hongchao, et al.
Published: (2024)
Surgical-VQLA++: Adversarial Contrastive Learning for Calibrated Robust Visual Question-Localized Answering in Robotic Surgery
by: Bai, Long, et al.
Published: (2024)
by: Bai, Long, et al.
Published: (2024)
Explainable AI for Automated User-specific Feedback in Surgical Skill Acquisition
by: Gomez, Catalina, et al.
Published: (2025)
by: Gomez, Catalina, et al.
Published: (2025)
Fast Reasoning Segmentation for Images and Videos
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
DualVision ArthroNav: Investigating Opportunities to Enhance Localization and Reconstruction in Image-based Arthroscopy Navigation via External Cameras
by: Shu, Hongchao, et al.
Published: (2025)
by: Shu, Hongchao, et al.
Published: (2025)
Humanoid Robots as First Assistants in Endoscopic Surgery
by: Cho, Sue Min, et al.
Published: (2026)
by: Cho, Sue Min, et al.
Published: (2026)
Counterfactual World Models via Digital Twin-conditioned Video Diffusion
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
Towards Robust Algorithms for Surgical Phase Recognition via Digital Twin Representation
by: Ding, Hao, et al.
Published: (2024)
by: Ding, Hao, et al.
Published: (2024)
Auto3DSeg for Brain Tumor Segmentation from 3D MRI in BraTS 2023 Challenge
by: Myronenko, Andriy, et al.
Published: (2025)
by: Myronenko, Andriy, et al.
Published: (2025)
Text-Driven Reasoning Video Editing via Reinforcement Learning on Digital Twin Representations
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
An Endoscopic Chisel: Intraoperative Imaging Carves 3D Anatomical Models
by: Mangulabnan, Jan Emily, et al.
Published: (2024)
by: Mangulabnan, Jan Emily, et al.
Published: (2024)
SDUM: A Scalable Deep Unrolled Model for Universal MRI Reconstruction
by: Wang, Puyang, et al.
Published: (2025)
by: Wang, Puyang, et al.
Published: (2025)
Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation
by: Li, Bohan, et al.
Published: (2026)
by: Li, Bohan, et al.
Published: (2026)
A Causal Framework for Aligning Image Quality Metrics and Deep Neural Network Robustness
by: Drenkow, Nathan, et al.
Published: (2025)
by: Drenkow, Nathan, et al.
Published: (2025)
Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
StraightTrack: Towards Mixed Reality Navigation System for Percutaneous K-wire Insertion
by: Zhang, Han, et al.
Published: (2024)
by: Zhang, Han, et al.
Published: (2024)
SAW-Bench: Learning Situated Awareness in the Real World
by: Li, Chuhan, et al.
Published: (2026)
by: Li, Chuhan, et al.
Published: (2026)
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
by: Zeng, Zhitao, et al.
Published: (2026)
by: Zeng, Zhitao, et al.
Published: (2026)
DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving
by: Shi, Chen, et al.
Published: (2026)
by: Shi, Chen, et al.
Published: (2026)
Promptable Counterfactual Diffusion Model for Unified Brain Tumor Segmentation and Generation with MRIs
by: Shen, Yiqing, et al.
Published: (2024)
by: Shen, Yiqing, et al.
Published: (2024)
Memorizing SAM: 3D Medical Segment Anything Model with Memorizing Transformer
by: Shao, Xinyuan, et al.
Published: (2024)
by: Shao, Xinyuan, et al.
Published: (2024)
Similar Items
-
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
by: Maksutova, Aiza, et al.
Published: (2026) -
Beyond Rigid AI: Towards Natural Human-Machine Symbiosis for Interoperative Surgical Assistance
by: Seenivasan, Lalithkumar, et al.
Published: (2025) -
Towards Controllable Video Synthesis of Routine and Rare OR Events
by: Schneider, Dominik, et al.
Published: (2026) -
From Generalization to Precision: Exploring SAM for Tool Segmentation in Surgical Environments
by: Oguine, Kanyifeechukwu J., et al.
Published: (2024) -
Online Reasoning Video Segmentation with Just-in-Time Digital Twins
by: Shen, Yiqing, et al.
Published: (2025)