Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Ochieng, Millicent, Gumma, Varun, Sitaram, Sunayana, Wang, Jindong, Chaudhary, Vishrav, Ronen, Keshet, Bali, Kalika, O'Neill, Jacki |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced Contexts
by: Ochieng, Millicent, et al.
Published: (2025)
by: Ochieng, Millicent, et al.
Published: (2025)
METAL: Towards Multilingual Meta-Evaluation
by: Hada, Rishav, et al.
Published: (2024)
by: Hada, Rishav, et al.
Published: (2024)
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
by: Ahuja, Sanchit, et al.
Published: (2023)
by: Ahuja, Sanchit, et al.
Published: (2023)
Contamination Report for Multilingual Benchmarks
by: Ahuja, Sanchit, et al.
Published: (2024)
by: Ahuja, Sanchit, et al.
Published: (2024)
HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
by: Gumma, Varun, et al.
Published: (2024)
by: Gumma, Varun, et al.
Published: (2024)
Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models
by: Gumma, Varun, et al.
Published: (2024)
by: Gumma, Varun, et al.
Published: (2024)
DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures
by: Seth, Agrima, et al.
Published: (2024)
by: Seth, Agrima, et al.
Published: (2024)
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
by: Hada, Rishav, et al.
Published: (2023)
by: Hada, Rishav, et al.
Published: (2023)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
by: Mukherjee, Sagnik, et al.
Published: (2024)
by: Mukherjee, Sagnik, et al.
Published: (2024)
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
by: Seth, Agrima, et al.
Published: (2025)
by: Seth, Agrima, et al.
Published: (2025)
PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
by: Watts, Ishaan, et al.
Published: (2024)
by: Watts, Ishaan, et al.
Published: (2024)
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
by: Jain, Prachi, et al.
Published: (2024)
by: Jain, Prachi, et al.
Published: (2024)
Dukawalla: Voice Interfaces for Small Businesses in Africa
by: Ankrah, Elizabeth, et al.
Published: (2025)
by: Ankrah, Elizabeth, et al.
Published: (2025)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
by: Kumar, Somnath, et al.
Published: (2024)
by: Kumar, Somnath, et al.
Published: (2024)
Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
by: Kumar, Somnath, et al.
Published: (2023)
by: Kumar, Somnath, et al.
Published: (2023)
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
by: Hamna, Hamna, et al.
Published: (2025)
by: Hamna, Hamna, et al.
Published: (2025)
CultureLLM: Incorporating Cultural Differences into Large Language Models
by: Li, Cheng, et al.
Published: (2024)
by: Li, Cheng, et al.
Published: (2024)
MunTTS: A Text-to-Speech System for Mundari
by: Gumma, Varun, et al.
Published: (2024)
by: Gumma, Varun, et al.
Published: (2024)
KAHANI: Culturally-Nuanced Visual Storytelling Tool for Non-Western Cultures
by: Hamna, et al.
Published: (2024)
by: Hamna, et al.
Published: (2024)
UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic Languages
by: Chitale, Pranjal A., et al.
Published: (2025)
by: Chitale, Pranjal A., et al.
Published: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
Improving Self Consistency in LLMs through Probabilistic Tokenization
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks
by: Schneider, Florian, et al.
Published: (2024)
by: Schneider, Florian, et al.
Published: (2024)
Towards Measuring and Modeling "Culture" in LLMs: A Survey
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
Akal Badi ya Bias: An Exploratory Study of Gender Bias in Hindi Language Technology
by: Hada, Rishav, et al.
Published: (2024)
by: Hada, Rishav, et al.
Published: (2024)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models
by: Aggarwal, Divyanshu, et al.
Published: (2024)
by: Aggarwal, Divyanshu, et al.
Published: (2024)
Giant Trees of Western America and the World, by Al Carder [Review]
by: O'Neill, Jim
Published: (2006)
by: O'Neill, Jim
Published: (2006)
Speech Representation Learning Revisited: The Necessity of Separate Learnable Parameters and Robust Data Augmentation
by: Yadav, Hemant, et al.
Published: (2024)
by: Yadav, Hemant, et al.
Published: (2024)
Analysing the Masked predictive coding training criterion for pre-training a Speech Representation Model
by: Yadav, Hemant, et al.
Published: (2023)
by: Yadav, Hemant, et al.
Published: (2023)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
by: Yadav, Hemant, et al.
Published: (2024)
by: Yadav, Hemant, et al.
Published: (2024)
JOOCI: a Framework for Learning Comprehensive Speech Representations
by: Yadav, Hemant, et al.
Published: (2024)
by: Yadav, Hemant, et al.
Published: (2024)
Multicultural Resources for Children: A Bibliography of Materials for Preschool Through Elementary School in the Areas of Black, Spanish-Speaking, Asian American, Native American, and Pacific Island Cultures.
by: Nichols, Margaret S., et al.
Published: (1977)
by: Nichols, Margaret S., et al.
Published: (1977)
DEPART: DEcomposing PARiTy across Multilingual LLMs
by: Uppadhyay, Manan, et al.
Published: (2026)
by: Uppadhyay, Manan, et al.
Published: (2026)
Online Numeric Data-Base Systems: A Resource for the Traditional Library.
by: Adams, Margaret O'Neill
Published: (1982)
by: Adams, Margaret O'Neill
Published: (1982)
World Perspective Case Descriptions on Educational Programs for Adults: Australia.
by: O'Neill, Barry, et al.
Published: (1989)
by: O'Neill, Barry, et al.
Published: (1989)
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
by: Aggarwal, Divyanshu, et al.
Published: (2024)
by: Aggarwal, Divyanshu, et al.
Published: (2024)
A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applications
by: Sitaram, Sunayana, et al.
Published: (2025)
by: Sitaram, Sunayana, et al.
Published: (2025)
Enhancing Low-Resource Minority Language Translation with LLMs and Retrieval-Augmented Generation for Cultural Nuances
by: Chang, Chen-Chi, et al.
Published: (2025)
by: Chang, Chen-Chi, et al.
Published: (2025)
sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting
by: Ahuja, Sanchit, et al.
Published: (2024)
by: Ahuja, Sanchit, et al.
Published: (2024)
Similar Items
-
Reasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced Contexts
by: Ochieng, Millicent, et al.
Published: (2025) -
METAL: Towards Multilingual Meta-Evaluation
by: Hada, Rishav, et al.
Published: (2024) -
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
by: Ahuja, Sanchit, et al.
Published: (2023) -
Contamination Report for Multilingual Benchmarks
by: Ahuja, Sanchit, et al.
Published: (2024) -
HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
by: Gumma, Varun, et al.
Published: (2024)