Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Nejadgholi, Isar, Kianpour, Masoud, Vishnubhotla, Krishnapriya, Molamohamadi, Maryam |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
by: Nejadgholi, Isar, et al.
Published: (2025)
by: Nejadgholi, Isar, et al.
Published: (2025)
The GPT-WritingPrompts Dataset: A Comparative Analysis of Character Portrayal in Short Stories
by: Huang, Xi Yu, et al.
Published: (2024)
by: Huang, Xi Yu, et al.
Published: (2024)
Tackling Social Bias against the Poor: A Dataset and Taxonomy on Aporophobia
by: Curto, Georgina, et al.
Published: (2025)
by: Curto, Georgina, et al.
Published: (2025)
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
by: Fraser, Kathleen C., et al.
Published: (2025)
by: Fraser, Kathleen C., et al.
Published: (2025)
The Emotion Dynamics of Literary Novels
by: Vishnubhotla, Krishnapriya, et al.
Published: (2024)
by: Vishnubhotla, Krishnapriya, et al.
Published: (2024)
WMT24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles
by: Dawkins, Hillary, et al.
Published: (2024)
by: Dawkins, Hillary, et al.
Published: (2024)
Gender-Neutral Machine Translation Strategies in Practice
by: Dawkins, Hillary, et al.
Published: (2025)
by: Dawkins, Hillary, et al.
Published: (2025)
Affect, Body, Cognition, Demographics, and Emotion: The ABCDE of Text Features for Computational Affective Science
by: Wahle, Jan Philip, et al.
Published: (2025)
by: Wahle, Jan Philip, et al.
Published: (2025)
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models
by: Ghanadian, Hamideh, et al.
Published: (2024)
by: Ghanadian, Hamideh, et al.
Published: (2024)
The crime of being poor
by: Curto, Georgina, et al.
Published: (2023)
by: Curto, Georgina, et al.
Published: (2023)
Projective Methods for Mitigating Gender Bias in Pre-trained Language Models
by: Dawkins, Hillary, et al.
Published: (2024)
by: Dawkins, Hillary, et al.
Published: (2024)
Human-Centered AI Applications for Canada's Immigration Settlement Sector
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
Emotion Granularity from Text: An Aggregate-Level Indicator of Mental Health
by: Vishnubhotla, Krishnapriya, et al.
Published: (2024)
by: Vishnubhotla, Krishnapriya, et al.
Published: (2024)
Social and Ethical Risks Posed by General-Purpose LLMs for Settling Newcomers in Canada
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
by: Guo, Rongchen, et al.
Published: (2024)
by: Guo, Rongchen, et al.
Published: (2024)
Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
by: Guo, Rongchen, et al.
Published: (2025)
by: Guo, Rongchen, et al.
Published: (2025)
From Perceived Effectiveness to Measured Impact: Identity-Aware Evaluation of Automated Counter-Stereotypes
by: Kiritchenko, Svetlana, et al.
Published: (2025)
by: Kiritchenko, Svetlana, et al.
Published: (2025)
Cultural Commonsense Knowledge for Intercultural Dialogues
by: Nguyen, Tuan-Phong, et al.
Published: (2024)
by: Nguyen, Tuan-Phong, et al.
Published: (2024)
ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs
by: Wang, Zhipin, et al.
Published: (2026)
by: Wang, Zhipin, et al.
Published: (2026)
Building a Custom Taxonomy of AI Skills and Tasks from the Ground Up with Job Postings
by: Meisenbacher, Stephen, et al.
Published: (2026)
by: Meisenbacher, Stephen, et al.
Published: (2026)
CultureForest: Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs
by: Ye, Yangfan, et al.
Published: (2026)
by: Ye, Yangfan, et al.
Published: (2026)
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding
by: Ye, Zhoutong, et al.
Published: (2025)
by: Ye, Zhoutong, et al.
Published: (2025)
Reference-Free Evaluation of Taxonomies
by: Wullschleger, Pascal, et al.
Published: (2025)
by: Wullschleger, Pascal, et al.
Published: (2025)
A Structured Framework for Evaluating and Enhancing Interpretive Capabilities of Multimodal LLMs in Culturally Situated Tasks
by: Yu, Haorui, et al.
Published: (2025)
by: Yu, Haorui, et al.
Published: (2025)
MentalBench: A DSM-Grounded Benchmark for Evaluating Psychiatric Diagnostic Capability of Large Language Models
by: Song, Hoyun, et al.
Published: (2026)
by: Song, Hoyun, et al.
Published: (2026)
Grounding Partially-Defined Events in Multimodal Data
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
Culturally Aware and Adapted NLP: A Taxonomy and a Survey of the State of the Art
by: Liu, Chen Cecilia, et al.
Published: (2024)
by: Liu, Chen Cecilia, et al.
Published: (2024)
GRAB: A Risk Taxonomy--Grounded Benchmark for Unsupervised Topic Discovery in Financial Disclosures
by: Li, Ying, et al.
Published: (2025)
by: Li, Ying, et al.
Published: (2025)
Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization
by: Sawant, Yash Ganpat
Published: (2026)
by: Sawant, Yash Ganpat
Published: (2026)
Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing
by: Dhar, Ruchira, et al.
Published: (2026)
by: Dhar, Ruchira, et al.
Published: (2026)
UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages
by: Abdullahi, Tassallah, et al.
Published: (2026)
by: Abdullahi, Tassallah, et al.
Published: (2026)
MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
by: Liu, Yinhong, et al.
Published: (2025)
by: Liu, Yinhong, et al.
Published: (2025)
Scaling Agentic Capabilities via Grounded Interaction Synthesis
by: Shi, Wenhang, et al.
Published: (2026)
by: Shi, Wenhang, et al.
Published: (2026)
GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians
by: Chen, Xiuyuan, et al.
Published: (2025)
by: Chen, Xiuyuan, et al.
Published: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools
by: Jung, Sam, et al.
Published: (2025)
by: Jung, Sam, et al.
Published: (2025)
Similar Items
-
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026) -
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
by: Nejadgholi, Isar, et al.
Published: (2025) -
The GPT-WritingPrompts Dataset: A Comparative Analysis of Character Portrayal in Short Stories
by: Huang, Xi Yu, et al.
Published: (2024) -
Tackling Social Bias against the Poor: A Dataset and Taxonomy on Aporophobia
by: Curto, Georgina, et al.
Published: (2025) -
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
by: Fraser, Kathleen C., et al.
Published: (2025)