Learning Surgical Robotic Manipulation with 3D Spatial Priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sheng, Yu, Wang, Lidian, Chu, Xiaomeng, Deng, Jiajun, Cheng, Min, Zhang, Yanyong, Hua, Bei, Li, Houqiang, Ji, Jianmin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914368140083200
author Sheng, Yu
Wang, Lidian
Chu, Xiaomeng
Deng, Jiajun
Cheng, Min
Zhang, Yanyong
Hua, Bei
Li, Houqiang
Ji, Jianmin
author_facet Sheng, Yu
Wang, Lidian
Chu, Xiaomeng
Deng, Jiajun
Cheng, Min
Zhang, Yanyong
Hua, Bei
Li, Houqiang
Ji, Jianmin
contents Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the surgical scene prior to manipulation, or enhance multi-view features by adding wrist-mounted cameras to supplement the default stereo endoscopes. However, both paradigms suffer from notable limitations: the former easily leads to error accumulation and prevents end-to-end optimization due to its multi-stage nature, while the latter is rarely adopted in clinical practice since wrist-mounted cameras can interfere with the motion of surgical robot arms. In this work, we introduce the Spatial Surgical Transformer (SST), an end-to-end visuomotor policy that empowers surgical robots with 3D spatial awareness by directly exploring 3D spatial cues embedded in endoscopic images. First, we build Surgical3D, a large-scale photorealistic dataset containing 30K stereo endoscopic image pairs with accurate 3D geometry, addressing the scarcity of 3D data in surgical scenes. Based on Surgical3D, we finetune a powerful geometric transformer to extract robust 3D latent representations from stereo endoscopes images. These representations are then seamlessly aligned with the robot's action space via a lightweight multi-level spatial feature connector (MSFC), all within an endoscope-centric coordinate frame. Extensive real-robot experiments demonstrate that SST achieves state-of-the-art performance and strong spatial generalization on complex surgical tasks such as knot tying and ex-vivo organ dissection, representing a significant step toward practical clinical deployment. The dataset and code will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03798
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning Surgical Robotic Manipulation with 3D Spatial Priors
Sheng, Yu
Wang, Lidian
Chu, Xiaomeng
Deng, Jiajun
Cheng, Min
Zhang, Yanyong
Hua, Bei
Li, Houqiang
Ji, Jianmin
Robotics
Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the surgical scene prior to manipulation, or enhance multi-view features by adding wrist-mounted cameras to supplement the default stereo endoscopes. However, both paradigms suffer from notable limitations: the former easily leads to error accumulation and prevents end-to-end optimization due to its multi-stage nature, while the latter is rarely adopted in clinical practice since wrist-mounted cameras can interfere with the motion of surgical robot arms. In this work, we introduce the Spatial Surgical Transformer (SST), an end-to-end visuomotor policy that empowers surgical robots with 3D spatial awareness by directly exploring 3D spatial cues embedded in endoscopic images. First, we build Surgical3D, a large-scale photorealistic dataset containing 30K stereo endoscopic image pairs with accurate 3D geometry, addressing the scarcity of 3D data in surgical scenes. Based on Surgical3D, we finetune a powerful geometric transformer to extract robust 3D latent representations from stereo endoscopes images. These representations are then seamlessly aligned with the robot's action space via a lightweight multi-level spatial feature connector (MSFC), all within an endoscope-centric coordinate frame. Extensive real-robot experiments demonstrate that SST achieves state-of-the-art performance and strong spatial generalization on complex surgical tasks such as knot tying and ex-vivo organ dissection, representing a significant step toward practical clinical deployment. The dataset and code will be released.
title Learning Surgical Robotic Manipulation with 3D Spatial Priors
topic Robotics
url https://arxiv.org/abs/2603.03798