Bengali Document Layout Analysis -- A YOLOV8 Based Ensembling Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahmed, Nazmus Sakib, Noor, Saad Sakib, Sikder, Ashraful Islam Shanto, Paul, Abhijit
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909814916907008
author Ahmed, Nazmus Sakib
Noor, Saad Sakib
Sikder, Ashraful Islam Shanto
Paul, Abhijit
author_facet Ahmed, Nazmus Sakib
Noor, Saad Sakib
Sikder, Ashraful Islam Shanto
Paul, Abhijit
contents This paper focuses on enhancing Bengali Document Layout Analysis (DLA) using the YOLOv8 model and innovative post-processing techniques. We tackle challenges unique to the complex Bengali script by employing data augmentation for model robustness. After meticulous validation set evaluation, we fine-tune our approach on the complete dataset, leading to a two-stage prediction strategy for accurate element segmentation. Our ensemble model, combined with post-processing, outperforms individual base architectures, addressing issues identified in the BaDLAD dataset. By leveraging this approach, we aim to advance Bengali document analysis, contributing to improved OCR and document comprehension and BaDLAD serves as a foundational resource for this endeavor, aiding future research in the field. Furthermore, our experiments provided key insights to incorporate new strategies into the established solution.
format Preprint
id arxiv_https___arxiv_org_abs_2309_00848
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Bengali Document Layout Analysis -- A YOLOV8 Based Ensembling Approach
Ahmed, Nazmus Sakib
Noor, Saad Sakib
Sikder, Ashraful Islam Shanto
Paul, Abhijit
Computer Vision and Pattern Recognition
Machine Learning
This paper focuses on enhancing Bengali Document Layout Analysis (DLA) using the YOLOv8 model and innovative post-processing techniques. We tackle challenges unique to the complex Bengali script by employing data augmentation for model robustness. After meticulous validation set evaluation, we fine-tune our approach on the complete dataset, leading to a two-stage prediction strategy for accurate element segmentation. Our ensemble model, combined with post-processing, outperforms individual base architectures, addressing issues identified in the BaDLAD dataset. By leveraging this approach, we aim to advance Bengali document analysis, contributing to improved OCR and document comprehension and BaDLAD serves as a foundational resource for this endeavor, aiding future research in the field. Furthermore, our experiments provided key insights to incorporate new strategies into the established solution.
title Bengali Document Layout Analysis -- A YOLOV8 Based Ensembling Approach
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2309.00848