SurgCUT3R: Surgical Scene-Aware Continuous Understanding of Temporal 3D Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Kaiyuan, Hong, Fangzhou, Elson, Daniel, Huang, Baoru |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SurgicalGS: Dynamic 3D Gaussian Splatting for Accurate Robotic-Assisted Surgical Scene Reconstruction
by: Chen, Jialei, et al.
Published: (2024)
by: Chen, Jialei, et al.
Published: (2024)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
by: Drago, Mauro Orazio, et al.
Published: (2025)
by: Drago, Mauro Orazio, et al.
Published: (2025)
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
by: Huang, Yiming, et al.
Published: (2025)
by: Huang, Yiming, et al.
Published: (2025)
3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
by: Zheng, Duo, et al.
Published: (2024)
by: Zheng, Duo, et al.
Published: (2024)
3D Scene Generation: A Survey
by: Wen, Beichen, et al.
Published: (2025)
by: Wen, Beichen, et al.
Published: (2025)
Surgical Visual Understanding (SurgVU) Dataset
by: Zia, Aneeq, et al.
Published: (2025)
by: Zia, Aneeq, et al.
Published: (2025)
R3DS: Reality-linked 3D Scenes for Panoramic Scene Understanding
by: Wu, Qirui, et al.
Published: (2024)
by: Wu, Qirui, et al.
Published: (2024)
Nested ResNet: A Vision-Based Method for Detecting the Sensing Area of a Drop-in Gamma Probe
by: Xu, Songyu, et al.
Published: (2024)
by: Xu, Songyu, et al.
Published: (2024)
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
High-fidelity Endoscopic Image Synthesis by Utilizing Depth-guided Neural Surfaces
by: Huang, Baoru, et al.
Published: (2024)
by: Huang, Baoru, et al.
Published: (2024)
SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature Alignment
by: Xu, Qi, et al.
Published: (2025)
by: Xu, Qi, et al.
Published: (2025)
Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding
by: Shi, Shuyao, et al.
Published: (2026)
by: Shi, Shuyao, et al.
Published: (2026)
SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
by: Choi, Tae-Min, et al.
Published: (2025)
by: Choi, Tae-Min, et al.
Published: (2025)
SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy
by: Li, Shi, et al.
Published: (2026)
by: Li, Shi, et al.
Published: (2026)
Laparoscopic Scene Analysis for Intraoperative Visualisation of Gamma Probe Signals in Minimally Invasive Cancer Surgery
by: Huang, Baoru
Published: (2025)
by: Huang, Baoru
Published: (2025)
3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
by: Wang, Xiaoye, et al.
Published: (2025)
by: Wang, Xiaoye, et al.
Published: (2025)
HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation
by: Biagini, Diego, et al.
Published: (2025)
by: Biagini, Diego, et al.
Published: (2025)
SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary
by: He, Jingyi, et al.
Published: (2026)
by: He, Jingyi, et al.
Published: (2026)
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos
by: Wu, Jinlin, et al.
Published: (2026)
by: Wu, Jinlin, et al.
Published: (2026)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
G-CUT3R: Guided 3D Reconstruction with Camera and Depth Prior Integration
by: Khafizov, Ramil, et al.
Published: (2025)
by: Khafizov, Ramil, et al.
Published: (2025)
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions
by: Cao, Yukang, et al.
Published: (2026)
by: Cao, Yukang, et al.
Published: (2026)
Enhancing Generalizability of Representation Learning for Data-Efficient 3D Scene Understanding
by: Wang, Yunsong, et al.
Published: (2024)
by: Wang, Yunsong, et al.
Published: (2024)
POMA-3D: The Point Map Way to 3D Scene Understanding
by: Mao, Ye, et al.
Published: (2025)
by: Mao, Ye, et al.
Published: (2025)
SurgTrack: CAD-Free 3D Tracking of Real-world Surgical Instruments
by: Guo, Wenwu, et al.
Published: (2024)
by: Guo, Wenwu, et al.
Published: (2024)
SurgFed: Language-guided Multi-Task Federated Learning for Surgical Video Understanding
by: Fang, Zheng, et al.
Published: (2026)
by: Fang, Zheng, et al.
Published: (2026)
Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025
by: Zia, Aneeq, et al.
Published: (2023)
by: Zia, Aneeq, et al.
Published: (2023)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Timeliness-Fidelity Tradeoff in 3D Scene Representations
by: Xu, Xiangmin, et al.
Published: (2024)
by: Xu, Xiangmin, et al.
Published: (2024)
StructLDM: Structured Latent Diffusion for 3D Human Generation
by: Hu, Tao, et al.
Published: (2024)
by: Hu, Tao, et al.
Published: (2024)
A Unified Framework for 3D Scene Understanding
by: Xu, Wei, et al.
Published: (2024)
by: Xu, Wei, et al.
Published: (2024)
SurgicalGaussian: Deformable 3D Gaussians for High-Fidelity Surgical Scene Reconstruction
by: Xie, Weixing, et al.
Published: (2024)
by: Xie, Weixing, et al.
Published: (2024)
EndoLRMGS: Complete Endoscopic Scene Reconstruction combining Large Reconstruction Modelling and Gaussian Splatting
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
HexPlane Representation for 3D Semantic Scene Understanding
by: Chen, Zeren, et al.
Published: (2025)
by: Chen, Zeren, et al.
Published: (2025)
SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention
by: Dhake, Shreyas C., et al.
Published: (2025)
by: Dhake, Shreyas C., et al.
Published: (2025)
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
by: Yu, Hanxun, et al.
Published: (2025)
by: Yu, Hanxun, et al.
Published: (2025)
AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
by: Vu, Nghia, et al.
Published: (2026)
by: Vu, Nghia, et al.
Published: (2026)
SurgAtt-Tracker: Online Surgical Attention Tracking via Temporal Proposal Reranking and Motion-Aware Refinement
by: Zhou, Rulin, et al.
Published: (2026)
by: Zhou, Rulin, et al.
Published: (2026)
Similar Items
-
SurgicalGS: Dynamic 3D Gaussian Splatting for Accurate Robotic-Assisted Surgical Scene Reconstruction
by: Chen, Jialei, et al.
Published: (2024) -
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
by: Drago, Mauro Orazio, et al.
Published: (2025) -
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
by: Huang, Yiming, et al.
Published: (2025) -
3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
by: Huang, Xiaohu, et al.
Published: (2025) -
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
by: Zheng, Duo, et al.
Published: (2024)