VS-Assistant: Versatile Surgery Assistant on the Demand of Surgeons
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zhen, Luo, Xingjian, Wu, Jinlin, Chan, Danny T. M., Lei, Zhen, Wang, Jinqiao, Ourselin, Sebastien, Liu, Hongbin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SurgPLAN++: Universal Surgical Phase Localization Network for Online and Offline Inference
by: Chen, Zhen, et al.
Published: (2024)
by: Chen, Zhen, et al.
Published: (2024)
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
ASI-Seg: Audio-Driven Surgical Instrument Segmentation with Surgeon Intention Understanding
by: Chen, Zhen, et al.
Published: (2024)
by: Chen, Zhen, et al.
Published: (2024)
SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement
by: Lei, Zeyu, et al.
Published: (2025)
by: Lei, Zeyu, et al.
Published: (2025)
EndoOmni: Zero-Shot Cross-Dataset Depth Estimation in Endoscopy by Robust Self-Learning from Noisy Labels
by: Tian, Qingyao, et al.
Published: (2024)
by: Tian, Qingyao, et al.
Published: (2024)
EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery
by: Wang, Guankun, et al.
Published: (2025)
by: Wang, Guankun, et al.
Published: (2025)
DaFoEs: Mixing Datasets towards the generalization of vision-state deep-learning Force Estimation in Minimally Invasive Robotic Surgery
by: Reyzabal, Mikel De Iturrate, et al.
Published: (2024)
by: Reyzabal, Mikel De Iturrate, et al.
Published: (2024)
SurgTrack: CAD-Free 3D Tracking of Real-world Surgical Instruments
by: Guo, Wenwu, et al.
Published: (2024)
by: Guo, Wenwu, et al.
Published: (2024)
Weakly Supervised YOLO Network for Surgical Instrument Localization in Endoscopic Videos
by: Wei, Rongfeng, et al.
Published: (2023)
by: Wei, Rongfeng, et al.
Published: (2023)
FoodLMM: A Versatile Food Assistant using Large Multi-modal Model
by: Yin, Yuehao, et al.
Published: (2023)
by: Yin, Yuehao, et al.
Published: (2023)
What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation
by: Chen, Zuyao, et al.
Published: (2024)
by: Chen, Zuyao, et al.
Published: (2024)
From Data to Modeling: Fully Open-vocabulary Scene Graph Generation
by: Chen, Zuyao, et al.
Published: (2025)
by: Chen, Zuyao, et al.
Published: (2025)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
Advancing Dense Endoscopic Reconstruction with Gaussian Splatting-driven Surface Normal-aware Tracking and Mapping
by: Huang, Yiming, et al.
Published: (2025)
by: Huang, Yiming, et al.
Published: (2025)
EndoMamba: An Efficient Foundation Model for Endoscopic Videos via Hierarchical Pre-training
by: Tian, Qingyao, et al.
Published: (2025)
by: Tian, Qingyao, et al.
Published: (2025)
SurgeMOD: Translating image-space tissue motions into vision-based surgical forces
by: Reyzabal, Mikel De Iturrate, et al.
Published: (2024)
by: Reyzabal, Mikel De Iturrate, et al.
Published: (2024)
TalkPhoto: A Versatile Training-Free Conversational Assistant for Intelligent Image Editing
by: Hu, Yujie, et al.
Published: (2026)
by: Hu, Yujie, et al.
Published: (2026)
Expanding Scene Graph Boundaries: Fully Open-vocabulary Scene Graph Generation via Visual-Concept Alignment and Retention
by: Chen, Zuyao, et al.
Published: (2023)
by: Chen, Zuyao, et al.
Published: (2023)
GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives
by: Chen, Zuyao, et al.
Published: (2023)
by: Chen, Zuyao, et al.
Published: (2023)
Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
by: Zhou, Lihua, et al.
Published: (2025)
by: Zhou, Lihua, et al.
Published: (2025)
SurgBox: Agent-Driven Operating Room Sandbox with Surgery Copilot
by: Wu, Jinlin, et al.
Published: (2024)
by: Wu, Jinlin, et al.
Published: (2024)
Anatomy-R1: Enhancing Anatomy Reasoning in Multimodal Large Language Models via Anatomical Similarity Curriculum and Group Diversity Augmentation
by: Song, Ziyang, et al.
Published: (2025)
by: Song, Ziyang, et al.
Published: (2025)
TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
by: Yan, Jiaqi, et al.
Published: (2025)
by: Yan, Jiaqi, et al.
Published: (2025)
Compile Scene Graphs with Reinforcement Learning
by: Chen, Zuyao, et al.
Published: (2025)
by: Chen, Zuyao, et al.
Published: (2025)
AgriDoctor: A Multimodal Intelligent Assistant for Agriculture
by: Zhang, Mingqing, et al.
Published: (2025)
by: Zhang, Mingqing, et al.
Published: (2025)
WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
EgoLife: Towards Egocentric Life Assistant
by: Yang, Jingkang, et al.
Published: (2025)
by: Yang, Jingkang, et al.
Published: (2025)
Endo-4DGX: Robust Endoscopic Scene Reconstruction and Illumination Correction with Gaussian Splatting
by: Huang, Yiming, et al.
Published: (2025)
by: Huang, Yiming, et al.
Published: (2025)
Consistent Assistant Domains Transformer for Source-free Domain Adaptation
by: Shao, Renrong, et al.
Published: (2025)
by: Shao, Renrong, et al.
Published: (2025)
Self-similarity Driven Scale-invariant Learning for Weakly Supervised Person Search
by: Wang, Benzhi, et al.
Published: (2023)
by: Wang, Benzhi, et al.
Published: (2023)
Co-Seg++: Mutual Prompt-Guided Collaborative Learning for Versatile Medical Segmentation
by: Xu, Qing, et al.
Published: (2025)
by: Xu, Qing, et al.
Published: (2025)
How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
DD-VNB: A Depth-based Dual-Loop Framework for Real-time Visually Navigated Bronchoscopy
by: Tian, Qingyao, et al.
Published: (2024)
by: Tian, Qingyao, et al.
Published: (2024)
HRVDA: High-Resolution Visual Document Assistant
by: Liu, Chaohu, et al.
Published: (2024)
by: Liu, Chaohu, et al.
Published: (2024)
SA-Person: Text-Based Person Retrieval with Scene-aware Re-ranking
by: Xu, Yingjia, et al.
Published: (2025)
by: Xu, Yingjia, et al.
Published: (2025)
Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation
by: Liang, Xusheng, et al.
Published: (2025)
by: Liang, Xusheng, et al.
Published: (2025)
Towards Realistic Hand-Object Interaction with Gravity-Field Based Diffusion Bridge
by: Xu, Miao, et al.
Published: (2025)
by: Xu, Miao, et al.
Published: (2025)
ArcSin: Adaptive ranged cosine Similarity injected noise for Language-Driven Visual Tasks
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
UniDA3D: A Unified Domain-Adaptive Framework for Multi-View 3D Object Detection
by: Wu, Hongjing, et al.
Published: (2026)
by: Wu, Hongjing, et al.
Published: (2026)
Customization Assistant for Text-to-image Generation
by: Zhou, Yufan, et al.
Published: (2023)
by: Zhou, Yufan, et al.
Published: (2023)
Similar Items
-
SurgPLAN++: Universal Surgical Phase Localization Network for Online and Offline Inference
by: Chen, Zhen, et al.
Published: (2024) -
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
by: Chen, Zhen, et al.
Published: (2025) -
ASI-Seg: Audio-Driven Surgical Instrument Segmentation with Surgeon Intention Understanding
by: Chen, Zhen, et al.
Published: (2024) -
SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement
by: Lei, Zeyu, et al.
Published: (2025) -
EndoOmni: Zero-Shot Cross-Dataset Depth Estimation in Endoscopy by Robust Self-Learning from Noisy Labels
by: Tian, Qingyao, et al.
Published: (2024)