mucAI at BAREC Shared Task 2025: Towards Uncertainty Aware Arabic Readability Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Abdou, Ahmed
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914046956011520
author Abdou, Ahmed
author_facet Abdou, Ahmed
contents We present a simple, model-agnostic post-processing technique for fine-grained Arabic readability classification in the BAREC 2025 Shared Task (19 ordinal levels). Our method applies conformal prediction to generate prediction sets with coverage guarantees, then computes weighted averages using softmax-renormalized probabilities over the conformal sets. This uncertainty-aware decoding improves Quadratic Weighted Kappa (QWK) by reducing high-penalty misclassifications to nearer levels. Our approach shows consistent QWK improvements of 1-3 points across different base models. In the strict track, our submission achieves QWK scores of 84.9\%(test) and 85.7\% (blind test) for sentence level, and 73.3\% for document level. For Arabic educational assessment, this enables human reviewers to focus on a handful of plausible levels, combining statistical guarantees with practical usability.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15485
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle mucAI at BAREC Shared Task 2025: Towards Uncertainty Aware Arabic Readability Assessment
Abdou, Ahmed
Computation and Language
Artificial Intelligence
We present a simple, model-agnostic post-processing technique for fine-grained Arabic readability classification in the BAREC 2025 Shared Task (19 ordinal levels). Our method applies conformal prediction to generate prediction sets with coverage guarantees, then computes weighted averages using softmax-renormalized probabilities over the conformal sets. This uncertainty-aware decoding improves Quadratic Weighted Kappa (QWK) by reducing high-penalty misclassifications to nearer levels. Our approach shows consistent QWK improvements of 1-3 points across different base models. In the strict track, our submission achieves QWK scores of 84.9\%(test) and 85.7\% (blind test) for sentence level, and 73.3\% for document level. For Arabic educational assessment, this enables human reviewers to focus on a handful of plausible levels, combining statistical guarantees with practical usability.
title mucAI at BAREC Shared Task 2025: Towards Uncertainty Aware Arabic Readability Assessment
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.15485