Benchmarking Sociolinguistic Diversity in Swahili NLP: A Taxonomy-Guided Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Oketch, Kezia, Lalor, John P., Abbasi, Ahmed |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the LLM Accessibility Divide? Performance, Fairness, and Cost of Closed versus Open LLMs for Automated Essay Scoring
by: Oketch, Kezia, et al.
Published: (2025)
by: Oketch, Kezia, et al.
Published: (2025)
Artificially Fluent: Swahili AI Performance Benchmarks Between English-Trained and Natively-Trained Datasets
by: Jaffer, Sophie, et al.
Published: (2025)
by: Jaffer, Sophie, et al.
Published: (2025)
PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins
by: Chen, Sihan, et al.
Published: (2025)
by: Chen, Sihan, et al.
Published: (2025)
Variation is the Norm: Embracing Sociolinguistics in NLP
by: Lutgen, Anne-Marie, et al.
Published: (2026)
by: Lutgen, Anne-Marie, et al.
Published: (2026)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
by: Jain, Shomik, et al.
Published: (2025)
by: Jain, Shomik, et al.
Published: (2025)
Bias A-head? Analyzing Bias in Transformer-Based Language Model Attention Heads
by: Yang, Yi, et al.
Published: (2023)
by: Yang, Yi, et al.
Published: (2023)
Understanding "Democratization" in NLP and ML Research
by: Subramonian, Arjun, et al.
Published: (2024)
by: Subramonian, Arjun, et al.
Published: (2024)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
by: Yang, Shujian, et al.
Published: (2025)
by: Yang, Shujian, et al.
Published: (2025)
Culture is Not Trivia: Sociocultural Theory for Cultural NLP
by: Zhou, Naitian, et al.
Published: (2025)
by: Zhou, Naitian, et al.
Published: (2025)
When a Nation Speaks: Machine Learning and NLP in People's Sentiment Analysis During Bangladesh's 2024 Mass Uprising
by: Alim, Md. Samiul, et al.
Published: (2025)
by: Alim, Md. Samiul, et al.
Published: (2025)
Applied Sociolinguistic AI for Community Development (ASA-CD): A New Scientific Paradigm for Linguistically-Grounded Social Intervention
by: Alam, S M Ruhul, et al.
Published: (2026)
by: Alam, S M Ruhul, et al.
Published: (2026)
NLP for Social Good: A Survey and Outlook of Challenges, Opportunities, and Responsible Deployment
by: Karamolegkou, Antonia, et al.
Published: (2025)
by: Karamolegkou, Antonia, et al.
Published: (2025)
Affective Computing in the Era of Large Language Models: A Survey from the NLP Perspective
by: Zhang, Yiqun, et al.
Published: (2024)
by: Zhang, Yiqun, et al.
Published: (2024)
Tackling Social Bias against the Poor: A Dataset and Taxonomy on Aporophobia
by: Curto, Georgina, et al.
Published: (2025)
by: Curto, Georgina, et al.
Published: (2025)
Automated Question Generation for Science Tests in Arabic Language Using NLP Techniques
by: Tami, Mohammad, et al.
Published: (2024)
by: Tami, Mohammad, et al.
Published: (2024)
Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
by: Gajewska, Ewelina, et al.
Published: (2025)
by: Gajewska, Ewelina, et al.
Published: (2025)
A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance
by: Melis, Matteo, et al.
Published: (2025)
by: Melis, Matteo, et al.
Published: (2025)
Uncovering Regulatory Affairs Complexity in Medical Products: A Qualitative Assessment Utilizing Open Coding and Natural Language Processing (NLP)
by: Han, Yu, et al.
Published: (2023)
by: Han, Yu, et al.
Published: (2023)
Impoverished Language Technology: The Lack of (Social) Class in NLP
by: Curry, Amanda Cercas, et al.
Published: (2024)
by: Curry, Amanda Cercas, et al.
Published: (2024)
A Taxonomy of Ambiguity Types for NLP
by: Li, Margaret Y., et al.
Published: (2024)
by: Li, Margaret Y., et al.
Published: (2024)
Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks
by: Alshehhi, Maitha, et al.
Published: (2025)
by: Alshehhi, Maitha, et al.
Published: (2025)
Ethical Concern Identification in NLP: A Corpus of ACL Anthology Ethics Statements
by: Karamolegkou, Antonia, et al.
Published: (2024)
by: Karamolegkou, Antonia, et al.
Published: (2024)
Domain-Independent Deception: A New Taxonomy and Linguistic Analysis
by: Verma, Rakesh M., et al.
Published: (2024)
by: Verma, Rakesh M., et al.
Published: (2024)
Addressing Both Statistical and Causal Gender Fairness in NLP Models
by: Chen, Hannah, et al.
Published: (2024)
by: Chen, Hannah, et al.
Published: (2024)
NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey
by: Goswami, Dhiman, et al.
Published: (2026)
by: Goswami, Dhiman, et al.
Published: (2026)
Stop! In the Name of Flaws: Disentangling Personal Names and Sociodemographic Attributes in NLP
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Decentralised Moderation for Interoperable Social Networks: A Conversation-based Approach for Pleroma and the Fediverse
by: Agarwal, Vibhor, et al.
Published: (2024)
by: Agarwal, Vibhor, et al.
Published: (2024)
SwaQuAD-24: QA Benchmark Dataset in Swahili
by: Kondoro, Alfred Malengo
Published: (2024)
by: Kondoro, Alfred Malengo
Published: (2024)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
by: Gomaa, Amr, et al.
Published: (2025)
by: Gomaa, Amr, et al.
Published: (2025)
NLP Cluster Analysis of Common Core State Standards and NAEP Item Specifications
by: Camilli, Gregory, et al.
Published: (2024)
by: Camilli, Gregory, et al.
Published: (2024)
Topic Classification of Case Law Using a Large Language Model and a New Taxonomy for UK Law: AI Insights into Summary Judgment
by: Sargeant, Holli, et al.
Published: (2024)
by: Sargeant, Holli, et al.
Published: (2024)
Human-Centric NLP or AI-Centric Illusion?: A Critical Investigation
by: Spencer, Piyapath T
Published: (2024)
by: Spencer, Piyapath T
Published: (2024)
Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AI
by: Liu, Houjiang, et al.
Published: (2023)
by: Liu, Houjiang, et al.
Published: (2023)
Towards Better Inclusivity: A Diverse Tweet Corpus of English Varieties
by: Pham, Nhi, et al.
Published: (2024)
by: Pham, Nhi, et al.
Published: (2024)
Multilingual Prompting for Improving LLM Generation Diversity
by: Wang, Qihan, et al.
Published: (2025)
by: Wang, Qihan, et al.
Published: (2025)
NLP Meets the World: Toward Improving Conversations With the Public About Natural Language Processing Research
by: Wilson, Shomir
Published: (2025)
by: Wilson, Shomir
Published: (2025)
From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP
by: Xu, Shanshan, et al.
Published: (2025)
by: Xu, Shanshan, et al.
Published: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications
by: Badhe, Sanket, et al.
Published: (2026)
by: Badhe, Sanket, et al.
Published: (2026)
A Taxonomy of Stereotype Content in Large Language Models
by: Nicolas, Gandalf, et al.
Published: (2024)
by: Nicolas, Gandalf, et al.
Published: (2024)
Similar Items
-
Bridging the LLM Accessibility Divide? Performance, Fairness, and Cost of Closed versus Open LLMs for Automated Essay Scoring
by: Oketch, Kezia, et al.
Published: (2025) -
Artificially Fluent: Swahili AI Performance Benchmarks Between English-Trained and Natively-Trained Datasets
by: Jaffer, Sophie, et al.
Published: (2025) -
PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins
by: Chen, Sihan, et al.
Published: (2025) -
Variation is the Norm: Embracing Sociolinguistics in NLP
by: Lutgen, Anne-Marie, et al.
Published: (2026) -
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
by: Jain, Shomik, et al.
Published: (2025)