Arabic Multi-turn Conversational Synthetic Dataset
Fuente:
Zenodo
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Recurso digital |
| Lingua: | arabo |
| Pubblicazione: |
Zenodo
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866902326472605696 |
|---|---|
| author | Misbah, Ahmed Farouk, Mohamed AbdulAzim, Mustafa |
| author_facet | Misbah, Ahmed Farouk, Mohamed AbdulAzim, Mustafa |
| contents | <p>A synthetic dataset of 43,316 conversations with mean conversation length of 14.038 turns (rounded to 3 decimal places), median of 12 turns, range of 5-111 turns, and a total of 608,052 utterances (where every turn is an utterance).</p> <p>Dataset is partitioned into training and test sets. An 80/20 split was adopted (34,653 training conversations / 8,663 test conversations).</p> <p>The synthetic data generation process systematically iterated over 93 topics and 151 countries, creating 14,043 unique topic-country combinations. The generation pipeline was configured to produce 5 conversations per combination. After rigorous processing and train/test split based on techniques to mitigate leakge risks, the end result was 43,316 conversations.</p> <p> </p> <p> </p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17855012 |
| institution | Zenodo |
| language | ara |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Arabic Multi-turn Conversational Synthetic Dataset Misbah, Ahmed Farouk, Mohamed AbdulAzim, Mustafa Arabic Conversations Chatbots <p>A synthetic dataset of 43,316 conversations with mean conversation length of 14.038 turns (rounded to 3 decimal places), median of 12 turns, range of 5-111 turns, and a total of 608,052 utterances (where every turn is an utterance).</p> <p>Dataset is partitioned into training and test sets. An 80/20 split was adopted (34,653 training conversations / 8,663 test conversations).</p> <p>The synthetic data generation process systematically iterated over 93 topics and 151 countries, creating 14,043 unique topic-country combinations. The generation pipeline was configured to produce 5 conversations per combination. After rigorous processing and train/test split based on techniques to mitigate leakge risks, the end result was 43,316 conversations.</p> <p> </p> <p> </p> |
| title | Arabic Multi-turn Conversational Synthetic Dataset |
| topic | Arabic Conversations Chatbots |
| url | https://doi.org/10.5281/zenodo.17855012 |