Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language
Fuente:
arXiv
Saved in:
| Main Authors: | Jimoh, Toheeb Aduramomi, De Wille, Tabea, Nikolov, Nikola S. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging Gaps in Natural Language Processing for Yorùbá: A Systematic Review of a Decade of Progress and Prospects
by: Jimoh, Toheeb Aduramomi, et al.
Published: (2025)
by: Jimoh, Toheeb Aduramomi, et al.
Published: (2025)
Attention-Based Deep Learning for Early Parkinson's Disease Detection with Tabular Biomedical Data
by: Oseni, Olamide Samuel, et al.
Published: (2026)
by: Oseni, Olamide Samuel, et al.
Published: (2026)
World model inspired sarcasm reasoning with large language model agents
by: Inoshita, Keito, et al.
Published: (2025)
by: Inoshita, Keito, et al.
Published: (2025)
InkubaLM: A small language model for low-resource African languages
by: Tonja, Atnafu Lambebo, et al.
Published: (2024)
by: Tonja, Atnafu Lambebo, et al.
Published: (2024)
Impact of emoji exclusion on the performance of Arabic sarcasm detection models
by: Aleryani, Ghalyah H., et al.
Published: (2024)
by: Aleryani, Ghalyah H., et al.
Published: (2024)
Assessing how hyperparameters impact Large Language Models' sarcasm detection performance
by: Gole, Montgomery, et al.
Published: (2025)
by: Gole, Montgomery, et al.
Published: (2025)
SemiAdapt and SemiLoRA: Efficient Domain Adaptation for Transformer-based Low-Resource Language Translation with a Case Study on Irish
by: McGiff, Josh, et al.
Published: (2025)
by: McGiff, Josh, et al.
Published: (2025)
Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo
by: Mbogho, Audrey, et al.
Published: (2025)
by: Mbogho, Audrey, et al.
Published: (2025)
Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review
by: McGiff, Josh, et al.
Published: (2025)
by: McGiff, Josh, et al.
Published: (2025)
Bridging the gap in online hate speech detection: a comparative analysis of BERT and traditional models for homophobic content identification on X/Twitter
by: McGiff, Josh, et al.
Published: (2024)
by: McGiff, Josh, et al.
Published: (2024)
synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier
by: Malik, Haq Nawaz, et al.
Published: (2026)
by: Malik, Haq Nawaz, et al.
Published: (2026)
How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language
by: Das, Dipto, et al.
Published: (2025)
by: Das, Dipto, et al.
Published: (2025)
Revealing the impact of synthetic native samples and multi-tasking strategies in Hindi-English code-mixed humour and sarcasm detection
by: Mazumder, Debajyoti, et al.
Published: (2024)
by: Mazumder, Debajyoti, et al.
Published: (2024)
Toxic language detection: a systematic review of Arabic datasets
by: Bensalem, Imene, et al.
Published: (2023)
by: Bensalem, Imene, et al.
Published: (2023)
Effective vocabulary expanding of multilingual language models for extremely low-resource languages
by: Zheng, Jianyu
Published: (2026)
by: Zheng, Jianyu
Published: (2026)
MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification
by: Alam, Kazi Samin Yasar, et al.
Published: (2026)
by: Alam, Kazi Samin Yasar, et al.
Published: (2026)
Can summarization approximate simplification? A gold standard comparison
by: Magnifico, Giacomo, et al.
Published: (2025)
by: Magnifico, Giacomo, et al.
Published: (2025)
Phonetically rich corpus construction for a low-resourced language
by: Amadeus, Marcellus, et al.
Published: (2024)
by: Amadeus, Marcellus, et al.
Published: (2024)
Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques
by: Xiong, Lang, et al.
Published: (2025)
by: Xiong, Lang, et al.
Published: (2025)
Leveraging LLMs for MT in Crisis Scenarios: a blueprint for low-resource languages
by: Lankford, Séamus, et al.
Published: (2024)
by: Lankford, Séamus, et al.
Published: (2024)
A comparison of pipelines for the translation of a low resource language based on transformers
by: Bonfanti, Chiara, et al.
Published: (2025)
by: Bonfanti, Chiara, et al.
Published: (2025)
Cross-lingual transfer of multilingual models on low resource African Languages
by: Thangaraj, Harish, et al.
Published: (2024)
by: Thangaraj, Harish, et al.
Published: (2024)
Multilingual jailbreaking of LLMs using low-resource languages
by: Marx, Dylan, et al.
Published: (2026)
by: Marx, Dylan, et al.
Published: (2026)
Prompt and circumstance: A word-by-word LLM prompting approach to interlinear glossing for low-resource languages
by: Elsner, Micha, et al.
Published: (2025)
by: Elsner, Micha, et al.
Published: (2025)
Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?
by: Dipto, Tawsif Tashwar, et al.
Published: (2025)
by: Dipto, Tawsif Tashwar, et al.
Published: (2025)
Development and bilingual evaluation of Japanese medical large language model within reasonably low computational resources
by: Sukeda, Issey
Published: (2024)
by: Sukeda, Issey
Published: (2024)
Natural language processing for African languages
by: Adelani, David Ifeoluwa
Published: (2025)
by: Adelani, David Ifeoluwa
Published: (2025)
Efficient Topic Extraction via Graph-Based Labeling: A Lightweight Alternative to Deep Models
by: Mekaoui, Salma, et al.
Published: (2025)
by: Mekaoui, Salma, et al.
Published: (2025)
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
by: Zhao, Tiancheng, et al.
Published: (2022)
by: Zhao, Tiancheng, et al.
Published: (2022)
A multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languages
by: Aliyu, Saminu Mohammad, et al.
Published: (2024)
by: Aliyu, Saminu Mohammad, et al.
Published: (2024)
FRACCO: A gold-standard annotated corpus of oncological entities with ICD-O-3.1 normalisation
by: Pignat, Johann, et al.
Published: (2025)
by: Pignat, Johann, et al.
Published: (2025)
Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages
by: Burda-Lassen, Olena
Published: (2024)
by: Burda-Lassen, Olena
Published: (2024)
UstanceBR: a social media language resource for stance prediction
by: Pereira, Camila, et al.
Published: (2023)
by: Pereira, Camila, et al.
Published: (2023)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
by: Paraskevopoulos, Georgios, et al.
Published: (2024)
by: Paraskevopoulos, Georgios, et al.
Published: (2024)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
by: Li, Kunning, et al.
Published: (2025)
by: Li, Kunning, et al.
Published: (2025)
NER- RoBERTa: Fine-Tuning RoBERTa for Named Entity Recognition (NER) within low-resource languages
by: Abdullah, Abdulhady Abas, et al.
Published: (2024)
by: Abdullah, Abdulhady Abas, et al.
Published: (2024)
Oddballness: universal anomaly detection with language models
by: Graliński, Filip, et al.
Published: (2024)
by: Graliński, Filip, et al.
Published: (2024)
Strategic resource allocation in memory encoding: An efficiency principle shaping language processing
by: Xu, Weijie, et al.
Published: (2025)
by: Xu, Weijie, et al.
Published: (2025)
Under-resourced studies of under-resourced languages: lemmatization and POS-tagging with LLM annotators for historical Armenian, Georgian, Greek and Syriac
by: Vidal-Gorène, Chahan, et al.
Published: (2026)
by: Vidal-Gorène, Chahan, et al.
Published: (2026)
Leveraging language models for summarizing mental state examinations: A comprehensive evaluation and dataset release
by: Sahu, Nilesh Kumar, et al.
Published: (2024)
by: Sahu, Nilesh Kumar, et al.
Published: (2024)
Similar Items
-
Bridging Gaps in Natural Language Processing for Yorùbá: A Systematic Review of a Decade of Progress and Prospects
by: Jimoh, Toheeb Aduramomi, et al.
Published: (2025) -
Attention-Based Deep Learning for Early Parkinson's Disease Detection with Tabular Biomedical Data
by: Oseni, Olamide Samuel, et al.
Published: (2026) -
World model inspired sarcasm reasoning with large language model agents
by: Inoshita, Keito, et al.
Published: (2025) -
InkubaLM: A small language model for low-resource African languages
by: Tonja, Atnafu Lambebo, et al.
Published: (2024) -
Impact of emoji exclusion on the performance of Arabic sarcasm detection models
by: Aleryani, Ghalyah H., et al.
Published: (2024)