Lexicon-Enriched Graph Modeling for Arabic Document Readability Prediction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Elchafei, Passant, Osama, Mayar, Rageh, Mohamed, Abuelkheir, Mervat
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911179385864192
author Elchafei, Passant
Osama, Mayar
Rageh, Mohamed
Abuelkheir, Mervat
author_facet Elchafei, Passant
Osama, Mayar
Rageh, Mohamed
Abuelkheir, Mervat
contents We present a graph-based approach enriched with lexicons to predict document-level readability in Arabic, developed as part of the Constrained Track of the BAREC Shared Task 2025. Our system models each document as a sentence-level graph, where nodes represent sentences and lemmas, and edges capture linguistic relationships such as lexical co-occurrence and class membership. Sentence nodes are enriched with features from the SAMER lexicon as well as contextual embeddings from the Arabic transformer model. The graph neural network (GNN) and transformer sentence encoder are trained as two independent branches, and their predictions are combined via late fusion at inference. For document-level prediction, sentence-level outputs are aggregated using max pooling to reflect the most difficult sentence. Experimental results show that this hybrid method outperforms standalone GNN or transformer branches across multiple readability metrics. Overall, the findings highlight that fusion offers advantages at the document level, but the GNN-only approach remains stronger for precise prediction of sentence-level readability.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22870
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lexicon-Enriched Graph Modeling for Arabic Document Readability Prediction
Elchafei, Passant
Osama, Mayar
Rageh, Mohamed
Abuelkheir, Mervat
Computation and Language
We present a graph-based approach enriched with lexicons to predict document-level readability in Arabic, developed as part of the Constrained Track of the BAREC Shared Task 2025. Our system models each document as a sentence-level graph, where nodes represent sentences and lemmas, and edges capture linguistic relationships such as lexical co-occurrence and class membership. Sentence nodes are enriched with features from the SAMER lexicon as well as contextual embeddings from the Arabic transformer model. The graph neural network (GNN) and transformer sentence encoder are trained as two independent branches, and their predictions are combined via late fusion at inference. For document-level prediction, sentence-level outputs are aggregated using max pooling to reflect the most difficult sentence. Experimental results show that this hybrid method outperforms standalone GNN or transformer branches across multiple readability metrics. Overall, the findings highlight that fusion offers advantages at the document level, but the GNN-only approach remains stronger for precise prediction of sentence-level readability.
title Lexicon-Enriched Graph Modeling for Arabic Document Readability Prediction
topic Computation and Language
url https://arxiv.org/abs/2509.22870