Saved in:
Bibliographic Details
Main Authors: Ghannam, Ahmad, Alharthi, Naif, Alasmary, Faris, Tabash, Kholood Al, Sadah, Shouq, Ghouti, Lahouari
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.24247
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911236588830720
author Ghannam, Ahmad
Alharthi, Naif
Alasmary, Faris
Tabash, Kholood Al
Sadah, Shouq
Ghouti, Lahouari
author_facet Ghannam, Ahmad
Alharthi, Naif
Alasmary, Faris
Tabash, Kholood Al
Sadah, Shouq
Ghouti, Lahouari
contents In this work, we tackle the Diacritic Restoration (DR) task for Arabic dialectal sentences using a multimodal approach that combines both textual and speech information. We propose a model that represents the text modality using an encoder extracted from our own pre-trained model named CATT. The speech component is handled by the encoder module of the OpenAI Whisper base model. Our solution is designed following two integration strategies. The former consists of fusing the speech tokens with the input at an early stage, where the 1500 frames of the audio segment are averaged over 10 consecutive frames, resulting in 150 speech tokens. To ensure embedding compatibility, these averaged tokens are processed through a linear projection layer prior to merging them with the text tokens. Contextual encoding is guaranteed by the CATT encoder module. The latter strategy relies on cross-attention, where text and speech embeddings are fused. The cross-attention output is then fed to the CATT classification head for token-level diacritic prediction. To further improve model robustness, we randomly deactivate the speech input during training, allowing the model to perform well with or without speech. Our experiments show that the proposed approach achieves a word error rate (WER) of 0.25 and a character error rate (CER) of 0.9 on the development set. On the test set, our model achieved WER and CER scores of 0.55 and 0.13, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24247
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Abjad AI at NADI 2025: CATT-Whisper: Multimodal Diacritic Restoration Using Text and Speech Representations
Ghannam, Ahmad
Alharthi, Naif
Alasmary, Faris
Tabash, Kholood Al
Sadah, Shouq
Ghouti, Lahouari
Computation and Language
In this work, we tackle the Diacritic Restoration (DR) task for Arabic dialectal sentences using a multimodal approach that combines both textual and speech information. We propose a model that represents the text modality using an encoder extracted from our own pre-trained model named CATT. The speech component is handled by the encoder module of the OpenAI Whisper base model. Our solution is designed following two integration strategies. The former consists of fusing the speech tokens with the input at an early stage, where the 1500 frames of the audio segment are averaged over 10 consecutive frames, resulting in 150 speech tokens. To ensure embedding compatibility, these averaged tokens are processed through a linear projection layer prior to merging them with the text tokens. Contextual encoding is guaranteed by the CATT encoder module. The latter strategy relies on cross-attention, where text and speech embeddings are fused. The cross-attention output is then fed to the CATT classification head for token-level diacritic prediction. To further improve model robustness, we randomly deactivate the speech input during training, allowing the model to perform well with or without speech. Our experiments show that the proposed approach achieves a word error rate (WER) of 0.25 and a character error rate (CER) of 0.9 on the development set. On the test set, our model achieved WER and CER scores of 0.55 and 0.13, respectively.
title Abjad AI at NADI 2025: CATT-Whisper: Multimodal Diacritic Restoration Using Text and Speech Representations
topic Computation and Language
url https://arxiv.org/abs/2510.24247