Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Guankun, Bai, Long, Nah, Wan Jun, Wang, Jie, Zhang, Zhaoxi, Chen, Zhen, Wu, Jinlin, Islam, Mobarakol, Liu, Hongbin, Ren, Hongliang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Surgical-VQLA++: Adversarial Contrastive Learning for Calibrated Robust Visual Question-Localized Answering in Robotic Surgery
by: Bai, Long, et al.
Published: (2024)
by: Bai, Long, et al.
Published: (2024)
SAM 2 in Robotic Surgery: An Empirical Evaluation for Robustness and Generalization in Surgical Video Segmentation
by: Yu, Jieming, et al.
Published: (2024)
by: Yu, Jieming, et al.
Published: (2024)
EndoDAC: Efficient Adapting Foundation Model for Self-Supervised Depth Estimation from Any Endoscopic Camera
by: Cui, Beilei, et al.
Published: (2024)
by: Cui, Beilei, et al.
Published: (2024)
Adapting SAM for Surgical Instrument Tracking and Segmentation in Endoscopic Submucosal Dissection Videos
by: Yu, Jieming, et al.
Published: (2024)
by: Yu, Jieming, et al.
Published: (2024)
Transferring Knowledge from High-Quality to Low-Quality MRI for Adult Glioma Diagnosis
by: Zhao, Yanguang, et al.
Published: (2024)
by: Zhao, Yanguang, et al.
Published: (2024)
EndoUIC: Promptable Diffusion Transformer for Unified Illumination Correction in Capsule Endoscopy
by: Bai, Long, et al.
Published: (2024)
by: Bai, Long, et al.
Published: (2024)
OSSAR: Towards Open-Set Surgical Activity Recognition in Robot-assisted Surgery
by: Bai, Long, et al.
Published: (2024)
by: Bai, Long, et al.
Published: (2024)
Surgical-DINO: Adapter Learning of Foundation Models for Depth Estimation in Endoscopic Surgery
by: Cui, Beilei, et al.
Published: (2024)
by: Cui, Beilei, et al.
Published: (2024)
EndoOOD: Uncertainty-aware Out-of-distribution Detection in Capsule Endoscopy Diagnosis
by: Tan, Qiaozhi, et al.
Published: (2024)
by: Tan, Qiaozhi, et al.
Published: (2024)
Privacy-Preserving Synthetic Continual Semantic Segmentation for Robotic Surgery
by: Xu, Mengya, et al.
Published: (2024)
by: Xu, Mengya, et al.
Published: (2024)
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
by: Huang, Yiming, et al.
Published: (2025)
by: Huang, Yiming, et al.
Published: (2025)
Weakly Supervised YOLO Network for Surgical Instrument Localization in Endoscopic Videos
by: Wei, Rongfeng, et al.
Published: (2023)
by: Wei, Rongfeng, et al.
Published: (2023)
Web-based Augmented Reality with Auto-Scaling and Real-Time Head Tracking towards Markerless Neurointerventional Preoperative Planning and Training of Head-mounted Robotic Needle Insertion
by: Ho, Hon Lung, et al.
Published: (2024)
by: Ho, Hon Lung, et al.
Published: (2024)
Illumination Histogram Consistency Metric for Quantitative Assessment of Video Sequences
by: Chen, Long, et al.
Published: (2024)
by: Chen, Long, et al.
Published: (2024)
EndoARSS: Adapting Spatially-Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery
by: Wang, Guankun, et al.
Published: (2025)
by: Wang, Guankun, et al.
Published: (2025)
EndoARSS: Adapting Spatially Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery
by: Guankun Wang, et al.
Published: (2025)
by: Guankun Wang, et al.
Published: (2025)
How can reasoning capability empower the AI copilot robot in endoscopic surgery
by: Wang, Guankun, et al.
Published: (2026)
by: Wang, Guankun, et al.
Published: (2026)
EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery
by: Wang, Guankun, et al.
Published: (2025)
by: Wang, Guankun, et al.
Published: (2025)
Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement
by: Bai, Long, et al.
Published: (2025)
by: Bai, Long, et al.
Published: (2025)
Unifying Image Processing as Visual Prompting Question Answering
by: Liu, Yihao, et al.
Published: (2023)
by: Liu, Yihao, et al.
Published: (2023)
LighTDiff: Surgical Endoscopic Image Low-Light Enhancement with T-Diffusion
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
by: Ma, Boyi, et al.
Published: (2025)
by: Ma, Boyi, et al.
Published: (2025)
PDZSeg: Adapting the Foundation Model for Dissection Zone Segmentation with Visual Prompts in Robot-assisted Endoscopic Submucosal Dissection
by: Xu, Mengya, et al.
Published: (2024)
by: Xu, Mengya, et al.
Published: (2024)
LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic Surgery
by: Du, Yuyang, et al.
Published: (2024)
by: Du, Yuyang, et al.
Published: (2024)
Efficient Bilinear Attention-based Fusion for Medical Visual Question Answering
by: Zhang, Zhilin, et al.
Published: (2024)
by: Zhang, Zhilin, et al.
Published: (2024)
Surgical-DeSAM: Decoupling SAM for Instrument Segmentation in Robotic Surgery
by: Sheng, Yuyang, et al.
Published: (2024)
by: Sheng, Yuyang, et al.
Published: (2024)
Goal-Oriented Semantic Communication for Wireless Visual Question Answering
by: Liu, Sige, et al.
Published: (2024)
by: Liu, Sige, et al.
Published: (2024)
SurgPLAN++: Universal Surgical Phase Localization Network for Online and Offline Inference
by: Chen, Zhen, et al.
Published: (2024)
by: Chen, Zhen, et al.
Published: (2024)
Toward Zero-Shot Learning for Visual Dehazing of Urological Surgical Robots
by: Wu, Renkai, et al.
Published: (2024)
by: Wu, Renkai, et al.
Published: (2024)
F2PASeg: Feature Fusion for Pituitary Anatomy Segmentation in Endoscopic Surgery
by: Chen, Lumin, et al.
Published: (2025)
by: Chen, Lumin, et al.
Published: (2025)
Visual Question Answering in Ophthalmology: A Progressive and Practical Perspective
by: Chen, Xiaolan, et al.
Published: (2024)
by: Chen, Xiaolan, et al.
Published: (2024)
ASI-Seg: Audio-Driven Surgical Instrument Segmentation with Surgeon Intention Understanding
by: Chen, Zhen, et al.
Published: (2024)
by: Chen, Zhen, et al.
Published: (2024)
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
by: Wang, Guankun, et al.
Published: (2025)
by: Wang, Guankun, et al.
Published: (2025)
CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
by: Wang, Guankun, et al.
Published: (2024)
by: Wang, Guankun, et al.
Published: (2024)
Geo-RepNet: Geometry-Aware Representation Learning for Surgical Phase Recognition in Endoscopic Submucosal Dissection
by: Tang, Rui, et al.
Published: (2025)
by: Tang, Rui, et al.
Published: (2025)
SAR-RAG: ATR Visual Question Answering by Semantic Search, Retrieval, and MLLM Generation
by: Ramirez, David F., et al.
Published: (2026)
by: Ramirez, David F., et al.
Published: (2026)
Learning to Efficiently Adapt Foundation Models for Self-Supervised Endoscopic 3D Scene Reconstruction from Any Cameras
by: Cui, Beilei, et al.
Published: (2025)
by: Cui, Beilei, et al.
Published: (2025)
Deep Learning for Surgical Instrument Recognition and Segmentation in Robotic-Assisted Surgeries: A Systematic Review
by: Ahmed, Fatimaelzahraa Ali, et al.
Published: (2024)
by: Ahmed, Fatimaelzahraa Ali, et al.
Published: (2024)
Towards a Large Language-Vision Question Answering Model for MSTAR Automatic Target Recognition
by: Ramirez, David F., et al.
Published: (2026)
by: Ramirez, David F., et al.
Published: (2026)
PitVQA: Image-grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery
by: He, Runlong, et al.
Published: (2024)
by: He, Runlong, et al.
Published: (2024)
Similar Items
-
Surgical-VQLA++: Adversarial Contrastive Learning for Calibrated Robust Visual Question-Localized Answering in Robotic Surgery
by: Bai, Long, et al.
Published: (2024) -
SAM 2 in Robotic Surgery: An Empirical Evaluation for Robustness and Generalization in Surgical Video Segmentation
by: Yu, Jieming, et al.
Published: (2024) -
EndoDAC: Efficient Adapting Foundation Model for Self-Supervised Depth Estimation from Any Endoscopic Camera
by: Cui, Beilei, et al.
Published: (2024) -
Adapting SAM for Surgical Instrument Tracking and Segmentation in Endoscopic Submucosal Dissection Videos
by: Yu, Jieming, et al.
Published: (2024) -
Transferring Knowledge from High-Quality to Low-Quality MRI for Adult Glioma Diagnosis
by: Zhao, Yanguang, et al.
Published: (2024)