Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce
Fuente:
arXiv
Saved in:
| Main Authors: | Ousidhoum, Nedjma, Beloucif, Meriem, Mohammad, Saif M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Annotating Dimensions of Social Perception in Text: A Sentence-Level Dataset of Warmth and Competence
by: Ayesh, Mutaz, et al.
Published: (2026)
by: Ayesh, Mutaz, et al.
Published: (2026)
Defining Boundaries: The Impact of Domain Specification on Cross-Language and Cross-Domain Transfer in Machine Translation
by: Shahnazaryan, Lia, et al.
Published: (2024)
by: Shahnazaryan, Lia, et al.
Published: (2024)
Causal Effects of Trigger Words in Social Media Discussions: A Large-Scale Case Study about UK Politics on Reddit
by: Antypas, Dimosthenis, et al.
Published: (2024)
by: Antypas, Dimosthenis, et al.
Published: (2024)
DIEKAE: Difference Injection for Efficient Knowledge Augmentation and Editing of Large Language Models
by: Galatolo, Alessio, et al.
Published: (2024)
by: Galatolo, Alessio, et al.
Published: (2024)
Words of Warmth: Trust and Sociability Norms for over 26k English Words
by: Mohammad, Saif M.
Published: (2025)
by: Mohammad, Saif M.
Published: (2025)
Social Good or Scientific Curiosity? Uncovering the Research Framing Behind NLP Artefacts
by: Chamoun, Eric, et al.
Published: (2025)
by: Chamoun, Eric, et al.
Published: (2025)
Insights from the ICLR Peer Review and Rebuttal Process
by: Kargaran, Amir Hossein, et al.
Published: (2025)
by: Kargaran, Amir Hossein, et al.
Published: (2025)
Visualising Policy-Reward Interplay to Inform Zeroth-Order Preference Optimisation of Large Language Models
by: Galatolo, Alessio, et al.
Published: (2025)
by: Galatolo, Alessio, et al.
Published: (2025)
SemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian Languages
by: Ousidhoum, Nedjma, et al.
Published: (2024)
by: Ousidhoum, Nedjma, et al.
Published: (2024)
Algorithmic Fragility and Persona Bias in LLM-Generated Autistic Communication
by: Rizvi, Naba, et al.
Published: (2026)
by: Rizvi, Naba, et al.
Published: (2026)
From Data Scarcity to Data Care: Reimagining Language Technologies for Serbian and other Low-Resource Languages
by: Ubois, Smiljana Antonijevic
Published: (2025)
by: Ubois, Smiljana Antonijevic
Published: (2025)
Mind the Gap: Pitfalls of LLM Alignment with Asian Public Opinion
by: Shankar, Hari, et al.
Published: (2026)
by: Shankar, Hari, et al.
Published: (2026)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
by: Sen, Indira, et al.
Published: (2023)
by: Sen, Indira, et al.
Published: (2023)
Better Call GPT, Comparing Large Language Models Against Lawyers
by: Martin, Lauren, et al.
Published: (2024)
by: Martin, Lauren, et al.
Published: (2024)
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
by: Anzenberg, Eitan, et al.
Published: (2025)
by: Anzenberg, Eitan, et al.
Published: (2025)
Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
Better Together: Quantifying the Benefits of AI-Assisted Recruitment
by: Aka, Ada, et al.
Published: (2025)
by: Aka, Ada, et al.
Published: (2025)
Mapping Violence: Developing an Extensive Framework to Build a Bangla Sectarian Expression Dataset from Social Media Interactions
by: Tasnim, Nazia, et al.
Published: (2024)
by: Tasnim, Nazia, et al.
Published: (2024)
Language Shift or Maintenance? An Intergenerational Study of the Tibetan Community in Saudi Arabia
by: Almoaily, Sumaiyah Turkistani Mohammad
Published: (2025)
by: Almoaily, Sumaiyah Turkistani Mohammad
Published: (2025)
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
by: Singh, Shrutika, et al.
Published: (2025)
by: Singh, Shrutika, et al.
Published: (2025)
Extracting O*NET Features from the NLx Corpus to Build Public Use Aggregate Labor Market Data
by: Meisenbacher, Stephen, et al.
Published: (2025)
by: Meisenbacher, Stephen, et al.
Published: (2025)
Beyond Keywords: Evaluating Large Language Model Classification of Nuanced Ableism
by: Rizvi, Naba, et al.
Published: (2025)
by: Rizvi, Naba, et al.
Published: (2025)
Uncovering Conspiratorial Narratives within Arabic Online Content
by: Mohdeb, Djamila, et al.
Published: (2025)
by: Mohdeb, Djamila, et al.
Published: (2025)
To Build Our Future, We Must Know Our Past: Contextualizing Paradigm Shifts in Natural Language Processing
by: Gururaja, Sireesh, et al.
Published: (2023)
by: Gururaja, Sireesh, et al.
Published: (2023)
Transformer-Based Low-Resource Language Translation: A Study on Standard Bengali to Sylheti
by: Oni, Mangsura Kabir, et al.
Published: (2025)
by: Oni, Mangsura Kabir, et al.
Published: (2025)
Towards Better Inclusivity: A Diverse Tweet Corpus of English Varieties
by: Pham, Nhi, et al.
Published: (2024)
by: Pham, Nhi, et al.
Published: (2024)
The Better Angels of Machine Personality: How Personality Relates to LLM Safety
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Multilinguality at the Edge: Developing Language Models for the Global South
by: Miranda, Lester James V., et al.
Published: (2026)
by: Miranda, Lester James V., et al.
Published: (2026)
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages
by: Ousidhoum, Nedjma, et al.
Published: (2024)
by: Ousidhoum, Nedjma, et al.
Published: (2024)
SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Detection
by: Muhammad, Shamsuddeen Hassan, et al.
Published: (2025)
by: Muhammad, Shamsuddeen Hassan, et al.
Published: (2025)
Automated Question Generation for Science Tests in Arabic Language Using NLP Techniques
by: Tami, Mohammad, et al.
Published: (2024)
by: Tami, Mohammad, et al.
Published: (2024)
AI-VERDE: A Gateway for Egalitarian Access to Large Language Model-Based Resources For Educational Institutions
by: Mithun, Paul, et al.
Published: (2025)
by: Mithun, Paul, et al.
Published: (2025)
AI Diffusion in Low Resource Language Countries
by: Misra, Amit, et al.
Published: (2025)
by: Misra, Amit, et al.
Published: (2025)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
by: Jin, Yiping, et al.
Published: (2024)
by: Jin, Yiping, et al.
Published: (2024)
Question-Answering (QA) Model for a Personalized Learning Assistant for Arabic Language
by: Sammoudi, Mohammad, et al.
Published: (2024)
by: Sammoudi, Mohammad, et al.
Published: (2024)
Development of Application-Specific Large Language Models to Facilitate Research Ethics Review
by: Mann, Sebastian Porsdam, et al.
Published: (2025)
by: Mann, Sebastian Porsdam, et al.
Published: (2025)
Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
by: Prama, Tabia Tanzin, et al.
Published: (2025)
by: Prama, Tabia Tanzin, et al.
Published: (2025)
Inverse Scaling: When Bigger Isn't Better
by: McKenzie, Ian R., et al.
Published: (2023)
by: McKenzie, Ian R., et al.
Published: (2023)
Towards Better Health Conversations: The Benefits of Context-seeking
by: Sayres, Rory, et al.
Published: (2025)
by: Sayres, Rory, et al.
Published: (2025)
Evaluating Machine Translation Datasets for Low-Web Data Languages: A Gendered Lens
by: Nigatu, Hellina Hailu, et al.
Published: (2025)
by: Nigatu, Hellina Hailu, et al.
Published: (2025)
Similar Items
-
Annotating Dimensions of Social Perception in Text: A Sentence-Level Dataset of Warmth and Competence
by: Ayesh, Mutaz, et al.
Published: (2026) -
Defining Boundaries: The Impact of Domain Specification on Cross-Language and Cross-Domain Transfer in Machine Translation
by: Shahnazaryan, Lia, et al.
Published: (2024) -
Causal Effects of Trigger Words in Social Media Discussions: A Large-Scale Case Study about UK Politics on Reddit
by: Antypas, Dimosthenis, et al.
Published: (2024) -
DIEKAE: Difference Injection for Efficient Knowledge Augmentation and Editing of Large Language Models
by: Galatolo, Alessio, et al.
Published: (2024) -
Words of Warmth: Trust and Sociability Norms for over 26k English Words
by: Mohammad, Saif M.
Published: (2025)