Arabic Text Diacritization In The Age Of Transfer Learning: Token Classification Is All You Need

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Skiredj, Abderrahman, Berrada, Ismail
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917563460485120
author Skiredj, Abderrahman
Berrada, Ismail
author_facet Skiredj, Abderrahman
Berrada, Ismail
contents Automatic diacritization of Arabic text involves adding diacritical marks (diacritics) to the text. This task poses a significant challenge with noteworthy implications for computational processing and comprehension. In this paper, we introduce PTCAD (Pre-FineTuned Token Classification for Arabic Diacritization, a novel two-phase approach for the Arabic Text Diacritization task. PTCAD comprises a pre-finetuning phase and a finetuning phase, treating Arabic Text Diacritization as a token classification task for pre-trained models. The effectiveness of PTCAD is demonstrated through evaluations on two benchmark datasets derived from the Tashkeela dataset, where it achieves state-of-the-art results, including a 20\% reduction in Word Error Rate (WER) compared to existing benchmarks and superior performance over GPT-4 in ATD tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04848
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Arabic Text Diacritization In The Age Of Transfer Learning: Token Classification Is All You Need
Skiredj, Abderrahman
Berrada, Ismail
Computation and Language
Automatic diacritization of Arabic text involves adding diacritical marks (diacritics) to the text. This task poses a significant challenge with noteworthy implications for computational processing and comprehension. In this paper, we introduce PTCAD (Pre-FineTuned Token Classification for Arabic Diacritization, a novel two-phase approach for the Arabic Text Diacritization task. PTCAD comprises a pre-finetuning phase and a finetuning phase, treating Arabic Text Diacritization as a token classification task for pre-trained models. The effectiveness of PTCAD is demonstrated through evaluations on two benchmark datasets derived from the Tashkeela dataset, where it achieves state-of-the-art results, including a 20\% reduction in Word Error Rate (WER) compared to existing benchmarks and superior performance over GPT-4 in ATD tasks.
title Arabic Text Diacritization In The Age Of Transfer Learning: Token Classification Is All You Need
topic Computation and Language
url https://arxiv.org/abs/2401.04848