EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ng, Chi Kit, Bai, Long, Wang, Guankun, Wang, Yupeng, Gao, Huxin, Yuan, Kun, Jin, Chenhan, Zeng, Tieyong, Ren, Hongliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911111904755712
author Ng, Chi Kit
Bai, Long
Wang, Guankun
Wang, Yupeng
Gao, Huxin
Yuan, Kun
Jin, Chenhan
Zeng, Tieyong
Ren, Hongliang
author_facet Ng, Chi Kit
Bai, Long
Wang, Guankun
Wang, Yupeng
Gao, Huxin
Yuan, Kun
Jin, Chenhan
Zeng, Tieyong
Ren, Hongliang
contents In endoscopic procedures, autonomous tracking of abnormal regions and following circumferential cutting markers can significantly reduce the cognitive burden on endoscopists. However, conventional model-based pipelines are fragile for each component (e.g., detection, motion planning) requires manual tuning and struggles to incorporate high-level endoscopic intent, leading to poor generalization across diverse scenes. Vision-Language-Action (VLA) models, which integrate visual perception, language grounding, and motion planning within an end-to-end framework, offer a promising alternative by semantically adapting to surgeon prompts without manual recalibration. Despite their potential, applying VLA models to robotic endoscopy presents unique challenges due to the complex and dynamic anatomical environments of the gastrointestinal (GI) tract. To address this, we introduce EndoVLA, designed specifically for continuum robots in GI interventions. Given endoscopic images and surgeon-issued tracking prompts, EndoVLA performs three core tasks: (1) polyp tracking, (2) delineation and following of abnormal mucosal regions, and (3) adherence to circular markers during circumferential cutting. To tackle data scarcity and domain shifts, we propose a dual-phase strategy comprising supervised fine-tuning on our EndoVLA-Motion dataset and reinforcement fine-tuning with task-aware rewards. Our approach significantly improves tracking performance in endoscopy and enables zero-shot generalization in diverse scenes and complex sequential tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15206
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy
Ng, Chi Kit
Bai, Long
Wang, Guankun
Wang, Yupeng
Gao, Huxin
Yuan, Kun
Jin, Chenhan
Zeng, Tieyong
Ren, Hongliang
Robotics
Artificial Intelligence
In endoscopic procedures, autonomous tracking of abnormal regions and following circumferential cutting markers can significantly reduce the cognitive burden on endoscopists. However, conventional model-based pipelines are fragile for each component (e.g., detection, motion planning) requires manual tuning and struggles to incorporate high-level endoscopic intent, leading to poor generalization across diverse scenes. Vision-Language-Action (VLA) models, which integrate visual perception, language grounding, and motion planning within an end-to-end framework, offer a promising alternative by semantically adapting to surgeon prompts without manual recalibration. Despite their potential, applying VLA models to robotic endoscopy presents unique challenges due to the complex and dynamic anatomical environments of the gastrointestinal (GI) tract. To address this, we introduce EndoVLA, designed specifically for continuum robots in GI interventions. Given endoscopic images and surgeon-issued tracking prompts, EndoVLA performs three core tasks: (1) polyp tracking, (2) delineation and following of abnormal mucosal regions, and (3) adherence to circular markers during circumferential cutting. To tackle data scarcity and domain shifts, we propose a dual-phase strategy comprising supervised fine-tuning on our EndoVLA-Motion dataset and reinforcement fine-tuning with task-aware rewards. Our approach significantly improves tracking performance in endoscopy and enables zero-shot generalization in diverse scenes and complex sequential tasks.
title EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2505.15206