EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy
Fuente:
arXiv
Saved in:
| Main Authors: | Ng, Chi Kit, Bai, Long, Wang, Guankun, Wang, Yupeng, Gao, Huxin, Yuan, Kun, Jin, Chenhan, Zeng, Tieyong, Ren, Hongliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TMR-VLA:Vision-Language-Action Model for Magnetic Motion Control of Tri-leg Silicone-based Soft Robot
by: Tang, Ruijie, et al.
Published: (2026)
by: Tang, Ruijie, et al.
Published: (2026)
EndoOOD: Uncertainty-aware Out-of-distribution Detection in Capsule Endoscopy Diagnosis
by: Tan, Qiaozhi, et al.
Published: (2024)
by: Tan, Qiaozhi, et al.
Published: (2024)
EndoARSS: Adapting Spatially-Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery
by: Wang, Guankun, et al.
Published: (2025)
by: Wang, Guankun, et al.
Published: (2025)
EndoARSS: Adapting Spatially Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery
by: Guankun Wang, et al.
Published: (2025)
by: Guankun Wang, et al.
Published: (2025)
Adapting SAM for Surgical Instrument Tracking and Segmentation in Endoscopic Submucosal Dissection Videos
by: Yu, Jieming, et al.
Published: (2024)
by: Yu, Jieming, et al.
Published: (2024)
GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning
by: Tang, Rui, et al.
Published: (2026)
by: Tang, Rui, et al.
Published: (2026)
Geo-RepNet: Geometry-Aware Representation Learning for Surgical Phase Recognition in Endoscopic Submucosal Dissection
by: Tang, Rui, et al.
Published: (2025)
by: Tang, Rui, et al.
Published: (2025)
How can reasoning capability empower the AI copilot robot in endoscopic surgery
by: Wang, Guankun, et al.
Published: (2026)
by: Wang, Guankun, et al.
Published: (2026)
Automatic Virtual‐to‐real Calibration and Dynamic Registration of Deformable Tissue for Endoscopic Submucosal Dissection
by: Yupeng Wang, et al.
Published: (2025)
by: Yupeng Wang, et al.
Published: (2025)
Jacobian Exploratory Dual-Phase Reinforcement Learning for Dynamic Endoluminal Navigation of Deformable Continuum Robots
by: Tian, Yu, et al.
Published: (2025)
by: Tian, Yu, et al.
Published: (2025)
Contact-Aided Navigation of Flexible Robotic Endoscope Using Deep Reinforcement Learning in Dynamic Stomach
by: Ng, Chi Kit, et al.
Published: (2025)
by: Ng, Chi Kit, et al.
Published: (2025)
CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
by: Wang, Guankun, et al.
Published: (2024)
by: Wang, Guankun, et al.
Published: (2024)
EndoDDC: Learning Sparse to Dense Reconstruction for Endoscopic Robotic Navigation via Diffusion Depth Completion
by: Lin, Yinheng, et al.
Published: (2026)
by: Lin, Yinheng, et al.
Published: (2026)
OSSAR: Towards Open-Set Surgical Activity Recognition in Robot-assisted Surgery
by: Bai, Long, et al.
Published: (2024)
by: Bai, Long, et al.
Published: (2024)
ETSM: Automating Dissection Trajectory Suggestion and Confidence Map-Based Safety Margin Prediction for Robot-assisted Endoscopic Submucosal Dissection
by: Xu, Mengya, et al.
Published: (2024)
by: Xu, Mengya, et al.
Published: (2024)
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
by: Ma, Boyi, et al.
Published: (2025)
by: Ma, Boyi, et al.
Published: (2025)
PDZSeg: Adapting the Foundation Model for Dissection Zone Segmentation with Visual Prompts in Robot-assisted Endoscopic Submucosal Dissection
by: Xu, Mengya, et al.
Published: (2024)
by: Xu, Mengya, et al.
Published: (2024)
Surgical-VQLA++: Adversarial Contrastive Learning for Calibrated Robust Visual Question-Localized Answering in Robotic Surgery
by: Bai, Long, et al.
Published: (2024)
by: Bai, Long, et al.
Published: (2024)
EndoUIC: Promptable Diffusion Transformer for Unified Illumination Correction in Capsule Endoscopy
by: Bai, Long, et al.
Published: (2024)
by: Bai, Long, et al.
Published: (2024)
Web-based Augmented Reality with Auto-Scaling and Real-Time Head Tracking towards Markerless Neurointerventional Preoperative Planning and Training of Head-mounted Robotic Needle Insertion
by: Ho, Hon Lung, et al.
Published: (2024)
by: Ho, Hon Lung, et al.
Published: (2024)
EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery
by: Wang, Guankun, et al.
Published: (2025)
by: Wang, Guankun, et al.
Published: (2025)
EndoDAC: Efficient Adapting Foundation Model for Self-Supervised Depth Estimation from Any Endoscopic Camera
by: Cui, Beilei, et al.
Published: (2024)
by: Cui, Beilei, et al.
Published: (2024)
Leveling3D: Leveling Up 3D Reconstruction with Feed-Forward 3D Gaussian Splatting and Geometry-Aware Generation
by: Huang, Yiming, et al.
Published: (2026)
by: Huang, Yiming, et al.
Published: (2026)
UnderwaterVLA: Dual-brain Vision-Language-Action architecture for Autonomous Underwater Navigation
by: Wang, Zhangyuan, et al.
Published: (2025)
by: Wang, Zhangyuan, et al.
Published: (2025)
EndoGSim: Physics-Aware 4D Dynamic Endoscopic Scene Simulations via MLLM-Guided Gaussian Splatting
by: Liu, Changjing, et al.
Published: (2026)
by: Liu, Changjing, et al.
Published: (2026)
Efficient Private SCO for Heavy-Tailed Data via Averaged Clipping
by: Jin, Chenhan, et al.
Published: (2022)
by: Jin, Chenhan, et al.
Published: (2022)
Endo-4DGS: Endoscopic Monocular Scene Reconstruction with 4D Gaussian Splatting
by: Huang, Yiming, et al.
Published: (2024)
by: Huang, Yiming, et al.
Published: (2024)
Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery
by: Wang, Guankun, et al.
Published: (2024)
by: Wang, Guankun, et al.
Published: (2024)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
CRL-VLA: Continual Vision-Language-Action Learning
by: Zeng, Qixin, et al.
Published: (2026)
by: Zeng, Qixin, et al.
Published: (2026)
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
by: Huang, Yiming, et al.
Published: (2025)
by: Huang, Yiming, et al.
Published: (2025)
Endo-TTAP: Robust Endoscopic Tissue Tracking via Multi-Facet Guided Attention and Hybrid Flow-point Supervision
by: Zhou, Rulin, et al.
Published: (2025)
by: Zhou, Rulin, et al.
Published: (2025)
EndoDINO: A Foundation Model for GI Endoscopy
by: Dermyer, Patrick, et al.
Published: (2025)
by: Dermyer, Patrick, et al.
Published: (2025)
RationalVLA: A Rational Vision-Language-Action Model with Dual System
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models
by: Zhang, Qiyao, et al.
Published: (2026)
by: Zhang, Qiyao, et al.
Published: (2026)
Endo-4DGX: Robust Endoscopic Scene Reconstruction and Illumination Correction with Gaussian Splatting
by: Huang, Yiming, et al.
Published: (2025)
by: Huang, Yiming, et al.
Published: (2025)
CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving
by: Arai, Hidehisa, et al.
Published: (2024)
by: Arai, Hidehisa, et al.
Published: (2024)
Learning to Adapt Foundation Model DINOv2 for Capsule Endoscopy Diagnosis
by: Zhang, Bowen, et al.
Published: (2024)
by: Zhang, Bowen, et al.
Published: (2024)
LighTDiff: Surgical Endoscopic Image Low-Light Enhancement with T-Diffusion
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
by: Luo, Yuechen, et al.
Published: (2026)
by: Luo, Yuechen, et al.
Published: (2026)
Similar Items
-
TMR-VLA:Vision-Language-Action Model for Magnetic Motion Control of Tri-leg Silicone-based Soft Robot
by: Tang, Ruijie, et al.
Published: (2026) -
EndoOOD: Uncertainty-aware Out-of-distribution Detection in Capsule Endoscopy Diagnosis
by: Tan, Qiaozhi, et al.
Published: (2024) -
EndoARSS: Adapting Spatially-Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery
by: Wang, Guankun, et al.
Published: (2025) -
EndoARSS: Adapting Spatially Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery
by: Guankun Wang, et al.
Published: (2025) -
Adapting SAM for Surgical Instrument Tracking and Segmentation in Endoscopic Submucosal Dissection Videos
by: Yu, Jieming, et al.
Published: (2024)