101 Billion Arabic Words Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Aloui, Manel, Chouikhi, Hasna, Chaabane, Ghaith, Kchaou, Haithem, Dhaouadi, Chehir |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GemmAr: Enhancing LLMs Through Arabic Instruction-Tuning
by: Chouikhi, Hasna, et al.
Published: (2024)
by: Chouikhi, Hasna, et al.
Published: (2024)
Noor-Ghateh: A Benchmark Dataset for Evaluating Arabic Word Segmenters in Hadith Domain
by: AlShuhayeb, Huda, et al.
Published: (2023)
by: AlShuhayeb, Huda, et al.
Published: (2023)
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
by: Darwish, Hazem, et al.
Published: (2024)
by: Darwish, Hazem, et al.
Published: (2024)
Advancing the Arabic WordNet: Elevating Content Quality
by: Freihat, Abed Alhakim, et al.
Published: (2024)
by: Freihat, Abed Alhakim, et al.
Published: (2024)
GASE: Generatively Augmented Sentence Encoding
by: Frank, Manuel, et al.
Published: (2024)
by: Frank, Manuel, et al.
Published: (2024)
CRaFT: An Explanation-Based Framework for Evaluating Cultural Reasoning in Multilingual Language Models
by: Hossain, Shehenaz, et al.
Published: (2025)
by: Hossain, Shehenaz, et al.
Published: (2025)
PTEB: Towards Robust Text Embedding Evaluation via Stochastic Paraphrasing at Evaluation Time with LLMs
by: Frank, Manuel, et al.
Published: (2025)
by: Frank, Manuel, et al.
Published: (2025)
Arabic Dataset for LLM Safeguard Evaluation
by: Ashraf, Yasser, et al.
Published: (2024)
by: Ashraf, Yasser, et al.
Published: (2024)
Gazelle: An Instruction Dataset for Arabic Writing Assistance
by: Magdy, Samar M., et al.
Published: (2024)
by: Magdy, Samar M., et al.
Published: (2024)
Testimole-Conversational: A 30-Billion-Word Italian Discussion Board Corpus (1996-2024) for Language Modeling and Sociolinguistic Research
by: Rinaldi, Matteo, et al.
Published: (2026)
by: Rinaldi, Matteo, et al.
Published: (2026)
Arabic Little STT: Arabic Children Speech Recognition Dataset
by: Alkadri, Mouhand, et al.
Published: (2025)
by: Alkadri, Mouhand, et al.
Published: (2025)
Arabic Multimodal Machine Learning: Datasets, Applications, Approaches, and Challenges
by: Haouhat, Abdelhamid, et al.
Published: (2025)
by: Haouhat, Abdelhamid, et al.
Published: (2025)
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
Proper Noun Diacritization for Arabic Wikipedia: A Benchmark Dataset
by: Bondok, Rawan, et al.
Published: (2025)
by: Bondok, Rawan, et al.
Published: (2025)
CIDAR: Culturally Relevant Instruction Dataset For Arabic
by: Alyafeai, Zaid, et al.
Published: (2024)
by: Alyafeai, Zaid, et al.
Published: (2024)
GLARE: Google Apps Arabic Reviews Dataset
by: AlGhamdi, Fatima, et al.
Published: (2024)
by: AlGhamdi, Fatima, et al.
Published: (2024)
Nullpointer at ArAIEval Shared Task: Arabic Propagandist Technique Detection with Token-to-Word Mapping in Sequence Tagging
by: Abir, Abrar, et al.
Published: (2024)
by: Abir, Abrar, et al.
Published: (2024)
VisCon-100K: Leveraging Contextual Web Data for Fine-tuning Vision Language Models
by: Kumar, Gokul Karthik, et al.
Published: (2025)
by: Kumar, Gokul Karthik, et al.
Published: (2025)
ArPoMeme: An Annotated Arabic Multimodal Dataset for Political Ideology and Polarization
by: Zaghouani, Wajdi, et al.
Published: (2026)
by: Zaghouani, Wajdi, et al.
Published: (2026)
LAILA: A Large Trait-Based Dataset for Arabic Automated Essay Scoring
by: Bashendy, May, et al.
Published: (2025)
by: Bashendy, May, et al.
Published: (2025)
EmoHopeSpeech: An Annotated Dataset of Emotions and Hope Speech in English and Arabic
by: Zaghouani, Wajdi, et al.
Published: (2025)
by: Zaghouani, Wajdi, et al.
Published: (2025)
Comparative Approaches to Sentiment Analysis Using Datasets in Major European and Arabic Languages
by: Krasitskii, Mikhail, et al.
Published: (2025)
by: Krasitskii, Mikhail, et al.
Published: (2025)
adaptMLLM: Fine-Tuning Multilingual Language Models on Low-Resource Languages with Integrated LLM Playgrounds
by: Lankford, Séamus, et al.
Published: (2024)
by: Lankford, Séamus, et al.
Published: (2024)
Design of an Open-Source Architecture for Neural Machine Translation
by: Lankford, Séamus, et al.
Published: (2024)
by: Lankford, Séamus, et al.
Published: (2024)
Human Evaluation of English--Irish Transformer-Based NMT
by: Lankford, Séamus, et al.
Published: (2024)
by: Lankford, Séamus, et al.
Published: (2024)
Machine Translation in the Covid domain: an English-Irish case study for LoResMT 2021
by: Lankford, Séamus, et al.
Published: (2024)
by: Lankford, Séamus, et al.
Published: (2024)
adaptNMT: an open-source, language-agnostic development environment for Neural Machine Translation
by: Lankford, Séamus, et al.
Published: (2024)
by: Lankford, Séamus, et al.
Published: (2024)
Transformers for Low-Resource Languages: Is Féidir Linn!
by: Lankford, Séamus, et al.
Published: (2024)
by: Lankford, Séamus, et al.
Published: (2024)
Towards Explainable Job Title Matching: Leveraging Semantic Textual Relatedness and Knowledge Graphs
by: Zadykian, Vadim, et al.
Published: (2025)
by: Zadykian, Vadim, et al.
Published: (2025)
ArabicaQA: A Comprehensive Dataset for Arabic Question Answering
by: Abdallah, Abdelrahman, et al.
Published: (2024)
by: Abdallah, Abdelrahman, et al.
Published: (2024)
Jawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking
by: Magdy, Samar M., et al.
Published: (2025)
by: Magdy, Samar M., et al.
Published: (2025)
AMuRD: Annotated Arabic-English Receipt Dataset for Key Information Extraction and Classification
by: Abdallah, Abdelrahman, et al.
Published: (2023)
by: Abdallah, Abdelrahman, et al.
Published: (2023)
Cohesion-6K: An Arabic Dataset for Analyzing Social Cohesion and Conflict in Online Discourse
by: Al-Athba, Aisha Ali, et al.
Published: (2026)
by: Al-Athba, Aisha Ali, et al.
Published: (2026)
MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark
by: Abu-Daoud, Mouath, et al.
Published: (2026)
by: Abu-Daoud, Mouath, et al.
Published: (2026)
PashtoCorp: A 1.25-Billion-Word Corpus, Evaluation Suite, and Reproducible Pipeline for Low-Resource Language Development
by: Rahman, Hanif
Published: (2026)
by: Rahman, Hanif
Published: (2026)
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
Advancing Dialectal Arabic to Modern Standard Arabic Machine Translation
by: Alabdullah, Abdullah, et al.
Published: (2025)
by: Alabdullah, Abdullah, et al.
Published: (2025)
CARMA: Comprehensive Automatically-annotated Reddit Mental Health Dataset for Arabic
by: Mankarious, Saad, et al.
Published: (2025)
by: Mankarious, Saad, et al.
Published: (2025)
Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
ADAB: Arabic Dataset for Automated Politeness Benchmarking -- A Large-Scale Resource for Computational Sociopragmatics
by: Al-Khalifa, Hend, et al.
Published: (2026)
by: Al-Khalifa, Hend, et al.
Published: (2026)
Similar Items
-
GemmAr: Enhancing LLMs Through Arabic Instruction-Tuning
by: Chouikhi, Hasna, et al.
Published: (2024) -
Noor-Ghateh: A Benchmark Dataset for Evaluating Arabic Word Segmenters in Hadith Domain
by: AlShuhayeb, Huda, et al.
Published: (2023) -
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
by: Darwish, Hazem, et al.
Published: (2024) -
Advancing the Arabic WordNet: Elevating Content Quality
by: Freihat, Abed Alhakim, et al.
Published: (2024) -
GASE: Generatively Augmented Sentence Encoding
by: Frank, Manuel, et al.
Published: (2024)