Gazelle: An Instruction Dataset for Arabic Writing Assistance
Fuente:
arXiv
Saved in:
| Main Authors: | Magdy, Samar M., Alwajih, Fakhraddin, Kwon, Sang Yun, Abdel-Salam, Reem, Abdul-Mageed, Muhammad |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking
by: Magdy, Samar M., et al.
Published: (2025)
by: Magdy, Samar M., et al.
Published: (2025)
Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic
by: Alwajih, Fakhraddin, et al.
Published: (2024)
by: Alwajih, Fakhraddin, et al.
Published: (2024)
Arabic Automatic Story Generation with Large Language Models
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation
by: Magdy, Samar M., et al.
Published: (2026)
by: Magdy, Samar M., et al.
Published: (2026)
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
Peacock: A Family of Arabic Multimodal Large Language Models and Benchmarks
by: Alwajih, Fakhraddin, et al.
Published: (2024)
by: Alwajih, Fakhraddin, et al.
Published: (2024)
PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
AfroScope: A Framework for Studying the Linguistic Landscape of Africa
by: Kwon, Sang Yun, et al.
Published: (2026)
by: Kwon, Sang Yun, et al.
Published: (2026)
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology
by: Sullivan, Peter, et al.
Published: (2026)
by: Sullivan, Peter, et al.
Published: (2026)
NADI 2024: The Fifth Nuanced Arabic Dialect Identification Shared Task
by: Abdul-Mageed, Muhammad, et al.
Published: (2024)
by: Abdul-Mageed, Muhammad, et al.
Published: (2024)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
by: Doan, Khai Duy, et al.
Published: (2024)
by: Doan, Khai Duy, et al.
Published: (2024)
Zero-Shot Context-Aware ASR for Diverse Arabic Varieties
by: Talafha, Bashar, et al.
Published: (2025)
by: Talafha, Bashar, et al.
Published: (2025)
Enhancing Health Mention Classification Performance: A Study on Advancements in Parameter Efficient Tuning
by: Abdel-Salam, Reem, et al.
Published: (2025)
by: Abdel-Salam, Reem, et al.
Published: (2025)
Voice of a Continent: Mapping Africa's Speech Technology Frontier
by: Elmadany, AbdelRahim, et al.
Published: (2025)
by: Elmadany, AbdelRahim, et al.
Published: (2025)
WojoodNER 2024: The Second Arabic Named Entity Recognition Shared Task
by: Jarrar, Mustafa, et al.
Published: (2024)
by: Jarrar, Mustafa, et al.
Published: (2024)
Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
Toucan: Many-to-Many Translation for 150 African Language Pairs
by: Elmadany, AbdelRahim, et al.
Published: (2024)
by: Elmadany, AbdelRahim, et al.
Published: (2024)
Cheetah: Natural Language Generation for 517 African Languages
by: Adebara, Ife, et al.
Published: (2024)
by: Adebara, Ife, et al.
Published: (2024)
Are LLMs Good Text Diacritizers? An Arabic and Yoruba Case Study
by: Toyin, Hawau Olamide, et al.
Published: (2025)
by: Toyin, Hawau Olamide, et al.
Published: (2025)
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
by: Toyin, Hawau Olamide, et al.
Published: (2025)
by: Toyin, Hawau Olamide, et al.
Published: (2025)
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition
by: Talafha, Bashar, et al.
Published: (2024)
by: Talafha, Bashar, et al.
Published: (2024)
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions
by: Wu, Minghao, et al.
Published: (2023)
by: Wu, Minghao, et al.
Published: (2023)
Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs
by: Mekki, Abdellah El, et al.
Published: (2026)
by: Mekki, Abdellah El, et al.
Published: (2026)
Effective Self-Mining of In-Context Examples for Unsupervised Machine Translation with LLMs
by: Mekki, Abdellah El, et al.
Published: (2024)
by: Mekki, Abdellah El, et al.
Published: (2024)
Estimating the Level of Dialectness Predicts Interannotator Agreement in Multi-dialect Arabic Datasets
by: Keleg, Amr, et al.
Published: (2024)
by: Keleg, Amr, et al.
Published: (2024)
CaresAI at BioCreative IX Track 1 -- LLM for Biomedical QA
by: Abdel-Salam, Reem, et al.
Published: (2025)
by: Abdel-Salam, Reem, et al.
Published: (2025)
Proper Noun Diacritization for Arabic Wikipedia: A Benchmark Dataset
by: Bondok, Rawan, et al.
Published: (2025)
by: Bondok, Rawan, et al.
Published: (2025)
NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
by: Talafha, Bashar, et al.
Published: (2025)
by: Talafha, Bashar, et al.
Published: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation
by: Waheed, Abdul, et al.
Published: (2024)
by: Waheed, Abdul, et al.
Published: (2024)
GLARE: Google Apps Arabic Reviews Dataset
by: AlGhamdi, Fatima, et al.
Published: (2024)
by: AlGhamdi, Fatima, et al.
Published: (2024)
On Barriers to Archival Audio Processing
by: Sullivan, Peter, et al.
Published: (2025)
by: Sullivan, Peter, et al.
Published: (2025)
Where Are We? Evaluating LLM Performance on African Languages
by: Adebara, Ife, et al.
Published: (2025)
by: Adebara, Ife, et al.
Published: (2025)
Do Biased Models Have Biased Thoughts?
by: Rajwal, Swati, et al.
Published: (2025)
by: Rajwal, Swati, et al.
Published: (2025)
Revisiting Common Assumptions about Arabic Dialects in NLP
by: Keleg, Amr, et al.
Published: (2025)
by: Keleg, Amr, et al.
Published: (2025)
CIDAR: Culturally Relevant Instruction Dataset For Arabic
by: Alyafeai, Zaid, et al.
Published: (2024)
by: Alyafeai, Zaid, et al.
Published: (2024)
ARWI: Arabic Write and Improve
by: Chirkunov, Kirill, et al.
Published: (2025)
by: Chirkunov, Kirill, et al.
Published: (2025)
LAILA: A Large Trait-Based Dataset for Arabic Automated Essay Scoring
by: Bashendy, May, et al.
Published: (2025)
by: Bashendy, May, et al.
Published: (2025)
Similar Items
-
Jawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking
by: Magdy, Samar M., et al.
Published: (2025) -
Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic
by: Alwajih, Fakhraddin, et al.
Published: (2024) -
Arabic Automatic Story Generation with Large Language Models
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024) -
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation
by: Magdy, Samar M., et al.
Published: (2026) -
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
by: Bhatia, Gagan, et al.
Published: (2024)