Arabic Multi-turn Conversational Synthetic Dataset

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autori principali: Misbah, Ahmed, Farouk, Mohamed, AbdulAzim, Mustafa
Natura: Recurso digital
Lingua:arabo
Pubblicazione: Zenodo 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866902326472605696
author Misbah, Ahmed
Farouk, Mohamed
AbdulAzim, Mustafa
author_facet Misbah, Ahmed
Farouk, Mohamed
AbdulAzim, Mustafa
contents <p>A synthetic dataset of 43,316 conversations with mean conversation length of 14.038 turns (rounded to 3 decimal places), median of 12 turns, range of 5-111 turns, and a total of 608,052 utterances (where every turn is an utterance).</p> <p>Dataset is partitioned into training and test sets. An 80/20 split was adopted (34,653 training conversations / 8,663 test conversations).</p> <p>The synthetic data generation process systematically iterated over 93 topics and 151 countries, creating 14,043 unique topic-country combinations. The generation pipeline was configured to produce 5 conversations per combination. After rigorous processing and train/test split based on techniques to mitigate leakge risks, the end result was 43,316 conversations.</p> <p> </p> <p> </p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17855012
institution Zenodo
language ara
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Arabic Multi-turn Conversational Synthetic Dataset
Misbah, Ahmed
Farouk, Mohamed
AbdulAzim, Mustafa
Arabic Conversations
Chatbots
<p>A synthetic dataset of 43,316 conversations with mean conversation length of 14.038 turns (rounded to 3 decimal places), median of 12 turns, range of 5-111 turns, and a total of 608,052 utterances (where every turn is an utterance).</p> <p>Dataset is partitioned into training and test sets. An 80/20 split was adopted (34,653 training conversations / 8,663 test conversations).</p> <p>The synthetic data generation process systematically iterated over 93 topics and 151 countries, creating 14,043 unique topic-country combinations. The generation pipeline was configured to produce 5 conversations per combination. After rigorous processing and train/test split based on techniques to mitigate leakge risks, the end result was 43,316 conversations.</p> <p> </p> <p> </p>
title Arabic Multi-turn Conversational Synthetic Dataset
topic Arabic Conversations
Chatbots
url https://doi.org/10.5281/zenodo.17855012