Capsule Vision Challenge 2024: Multi-Class Abnormality Classification for Video Capsule Endoscopy

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bansal, Aakarsh, Singla, Bhuvanesh, Wankhade, Raajan Rajesh, Patil, Nagamma
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910682077724672
author Bansal, Aakarsh
Singla, Bhuvanesh
Wankhade, Raajan Rajesh
Patil, Nagamma
author_facet Bansal, Aakarsh
Singla, Bhuvanesh
Wankhade, Raajan Rajesh
Patil, Nagamma
contents This study presents an approach to developing a model for classifying abnormalities in video capsule endoscopy (VCE) frames. Given the challenges of data imbalance, we implemented a tiered augmentation strategy using the albumentations library to enhance minority class representation. Additionally, we addressed learning complexities by progressively structuring training tasks, allowing the model to differentiate between normal and abnormal cases and then gradually adding more specific classes based on data availability. Our pipeline, developed in PyTorch, employs a flexible architecture enabling seamless adjustments to classification complexity. We tested our approach using ResNet50 and a custom ViT-CNN hybrid model, with training conducted on the Kaggle platform. This work demonstrates a scalable approach to abnormality classification in VCE.
format Preprint
id arxiv_https___arxiv_org_abs_2411_01479
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Capsule Vision Challenge 2024: Multi-Class Abnormality Classification for Video Capsule Endoscopy
Bansal, Aakarsh
Singla, Bhuvanesh
Wankhade, Raajan Rajesh
Patil, Nagamma
Computer Vision and Pattern Recognition
Artificial Intelligence
This study presents an approach to developing a model for classifying abnormalities in video capsule endoscopy (VCE) frames. Given the challenges of data imbalance, we implemented a tiered augmentation strategy using the albumentations library to enhance minority class representation. Additionally, we addressed learning complexities by progressively structuring training tasks, allowing the model to differentiate between normal and abnormal cases and then gradually adding more specific classes based on data availability. Our pipeline, developed in PyTorch, employs a flexible architecture enabling seamless adjustments to classification complexity. We tested our approach using ResNet50 and a custom ViT-CNN hybrid model, with training conducted on the Kaggle platform. This work demonstrates a scalable approach to abnormality classification in VCE.
title Capsule Vision Challenge 2024: Multi-Class Abnormality Classification for Video Capsule Endoscopy
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.01479