EndoGen: Conditional Autoregressive Endoscopic Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xinyu, Liu, Hengyu, Wang, Cheng, Liu, Tianming, Yuan, Yixuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915406117076992
author Liu, Xinyu
Liu, Hengyu
Wang, Cheng
Liu, Tianming
Yuan, Yixuan
author_facet Liu, Xinyu
Liu, Hengyu
Wang, Cheng
Liu, Tianming
Yuan, Yixuan
contents Endoscopic video generation is crucial for advancing medical imaging and enhancing diagnostic capabilities. However, prior efforts in this field have either focused on static images, lacking the dynamic context required for practical applications, or have relied on unconditional generation that fails to provide meaningful references for clinicians. Therefore, in this paper, we propose the first conditional endoscopic video generation framework, namely EndoGen. Specifically, we build an autoregressive model with a tailored Spatiotemporal Grid-Frame Patterning (SGP) strategy. It reformulates the learning of generating multiple frames as a grid-based image generation pattern, which effectively capitalizes the inherent global dependency modeling capabilities of autoregressive architectures. Furthermore, we propose a Semantic-Aware Token Masking (SAT) mechanism, which enhances the model's ability to produce rich and diverse content by selectively focusing on semantically meaningful regions during the generation process. Through extensive experiments, we demonstrate the effectiveness of our framework in generating high-quality, conditionally guided endoscopic content, and improves the performance of downstream task of polyp segmentation. Code released at https://www.github.com/CUHK-AIM-Group/EndoGen.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17388
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EndoGen: Conditional Autoregressive Endoscopic Video Generation
Liu, Xinyu
Liu, Hengyu
Wang, Cheng
Liu, Tianming
Yuan, Yixuan
Computer Vision and Pattern Recognition
Image and Video Processing
Endoscopic video generation is crucial for advancing medical imaging and enhancing diagnostic capabilities. However, prior efforts in this field have either focused on static images, lacking the dynamic context required for practical applications, or have relied on unconditional generation that fails to provide meaningful references for clinicians. Therefore, in this paper, we propose the first conditional endoscopic video generation framework, namely EndoGen. Specifically, we build an autoregressive model with a tailored Spatiotemporal Grid-Frame Patterning (SGP) strategy. It reformulates the learning of generating multiple frames as a grid-based image generation pattern, which effectively capitalizes the inherent global dependency modeling capabilities of autoregressive architectures. Furthermore, we propose a Semantic-Aware Token Masking (SAT) mechanism, which enhances the model's ability to produce rich and diverse content by selectively focusing on semantically meaningful regions during the generation process. Through extensive experiments, we demonstrate the effectiveness of our framework in generating high-quality, conditionally guided endoscopic content, and improves the performance of downstream task of polyp segmentation. Code released at https://www.github.com/CUHK-AIM-Group/EndoGen.
title EndoGen: Conditional Autoregressive Endoscopic Video Generation
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2507.17388