EmoAra: Emotion-Preserving English Speech Transcription and Cross-Lingual Translation with Arabic Text-to-Speech

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hassan, Besher, Alsarraj, Ibrahim, Hasan, Musaab, Melhim, Yousef, Fadi, Shahem, Sultan, Shahem
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915767284400128
author Hassan, Besher
Alsarraj, Ibrahim
Hasan, Musaab
Melhim, Yousef
Fadi, Shahem
Sultan, Shahem
author_facet Hassan, Besher
Alsarraj, Ibrahim
Hasan, Musaab
Melhim, Yousef
Fadi, Shahem
Sultan, Shahem
contents This work presents EmoAra, an end-to-end emotion-preserving pipeline for cross-lingual spoken communication, motivated by banking customer service where emotional context affects service quality. EmoAra integrates Speech Emotion Recognition, Automatic Speech Recognition, Machine Translation, and Text-to-Speech to process English speech and deliver an Arabic spoken output while retaining emotional nuance. The system uses a CNN-based emotion classifier, Whisper for English transcription, a fine-tuned MarianMT model for English-to-Arabic translation, and MMS-TTS-Ara for Arabic speech synthesis. Experiments report an F1-score of 94% for emotion classification, translation performance of BLEU 56 and BERTScore F1 88.7%, and an average human evaluation score of 81% on banking-domain translations. The implementation and resources are available at the accompanying GitHub repository.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01170
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EmoAra: Emotion-Preserving English Speech Transcription and Cross-Lingual Translation with Arabic Text-to-Speech
Hassan, Besher
Alsarraj, Ibrahim
Hasan, Musaab
Melhim, Yousef
Fadi, Shahem
Sultan, Shahem
Computation and Language
This work presents EmoAra, an end-to-end emotion-preserving pipeline for cross-lingual spoken communication, motivated by banking customer service where emotional context affects service quality. EmoAra integrates Speech Emotion Recognition, Automatic Speech Recognition, Machine Translation, and Text-to-Speech to process English speech and deliver an Arabic spoken output while retaining emotional nuance. The system uses a CNN-based emotion classifier, Whisper for English transcription, a fine-tuned MarianMT model for English-to-Arabic translation, and MMS-TTS-Ara for Arabic speech synthesis. Experiments report an F1-score of 94% for emotion classification, translation performance of BLEU 56 and BERTScore F1 88.7%, and an average human evaluation score of 81% on banking-domain translations. The implementation and resources are available at the accompanying GitHub repository.
title EmoAra: Emotion-Preserving English Speech Transcription and Cross-Lingual Translation with Arabic Text-to-Speech
topic Computation and Language
url https://arxiv.org/abs/2602.01170