Seeing More with Less: Video Capsule Endoscopy with Multi-Task Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Werner, Julia, Bause, Oliver, Oexle, Julius, Floch, Maxime Le, Brinkmann, Franz, Hampe, Jochen, Bringmann, Oliver
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915738759987200
author Werner, Julia
Bause, Oliver
Oexle, Julius
Floch, Maxime Le
Brinkmann, Franz
Hampe, Jochen
Bringmann, Oliver
author_facet Werner, Julia
Bause, Oliver
Oexle, Julius
Floch, Maxime Le
Brinkmann, Franz
Hampe, Jochen
Bringmann, Oliver
contents Video capsule endoscopy has become increasingly important for investigating the small intestine within the gastrointestinal tract. However, a persistent challenge remains the short battery lifetime of such compact sensor edge devices. Integrating artificial intelligence can help overcome this limitation by enabling intelligent real-time decision-making, thereby reducing the energy consumption and prolonging the battery life. However, this remains challenging due to data sparsity and the limited resources of the device restricting the overall model size. In this work, we introduce a multi-task neural network that combines the functionalities of precise self-localization within the gastrointestinal tract with the ability to detect anomalies in the small intestine within a single model. Throughout the development process, we consistently restricted the total number of parameters to ensure the feasibility to deploy such model in a small capsule. We report the first multi-task results using the recently published Galar dataset, integrating established multi-task methods and Viterbi decoding for subsequent time-series analysis. This outperforms current single-task models and represents a significant advance in AI-based approaches in this field. Our model achieves an accuracy of 93.63% on the localization task and an accuracy of 87.48% on the anomaly detection task. The approach requires only 1 million parameters while surpassing the current baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23479
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seeing More with Less: Video Capsule Endoscopy with Multi-Task Learning
Werner, Julia
Bause, Oliver
Oexle, Julius
Floch, Maxime Le
Brinkmann, Franz
Hampe, Jochen
Bringmann, Oliver
Computer Vision and Pattern Recognition
Video capsule endoscopy has become increasingly important for investigating the small intestine within the gastrointestinal tract. However, a persistent challenge remains the short battery lifetime of such compact sensor edge devices. Integrating artificial intelligence can help overcome this limitation by enabling intelligent real-time decision-making, thereby reducing the energy consumption and prolonging the battery life. However, this remains challenging due to data sparsity and the limited resources of the device restricting the overall model size. In this work, we introduce a multi-task neural network that combines the functionalities of precise self-localization within the gastrointestinal tract with the ability to detect anomalies in the small intestine within a single model. Throughout the development process, we consistently restricted the total number of parameters to ensure the feasibility to deploy such model in a small capsule. We report the first multi-task results using the recently published Galar dataset, integrating established multi-task methods and Viterbi decoding for subsequent time-series analysis. This outperforms current single-task models and represents a significant advance in AI-based approaches in this field. Our model achieves an accuracy of 93.63% on the localization task and an accuracy of 87.48% on the anomaly detection task. The approach requires only 1 million parameters while surpassing the current baselines.
title Seeing More with Less: Video Capsule Endoscopy with Multi-Task Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.23479