SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ousidhoum, Nedjma, Myung, Junho, Perez-Almendros, Carla, Jin, Jiho, Keleg, Amr, Beloucif, Meriem, Zhou, Yi, Agerri, Rodrigo, Araujo, Vladimir, Baes, Naomi, Barry, James, Boisson, Joanne, Chen, Nancy F., de Kock, Christine, Edwards, Aleksandra, de Landa, Joseba Fernandez, Imam, Mohamed Fazli, Hakami, Huda, Hsieh, Shu-Kai, Imperial, Joseph Marvin, Lee, Roy Ka-Wei, Liu, Zhengyuan, Lyu, Chenyang, Samih, Younes, Sjons, Johan, Tan, Bryan, Ushio, Asahi, Zheng, Weihua, Oh, Alice, Camacho-Collados, Jose
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914528078331904
author Ousidhoum, Nedjma
Myung, Junho
Perez-Almendros, Carla
Jin, Jiho
Keleg, Amr
Beloucif, Meriem
Zhou, Yi
Agerri, Rodrigo
Araujo, Vladimir
Baes, Naomi
Barry, James
Boisson, Joanne
Chen, Nancy F.
de Kock, Christine
Edwards, Aleksandra
de Landa, Joseba Fernandez
Imam, Mohamed Fazli
Hakami, Huda
Hsieh, Shu-Kai
Imperial, Joseph Marvin
Lee, Roy Ka-Wei
Liu, Zhengyuan
Lyu, Chenyang
Samih, Younes
Sjons, Johan
Tan, Bryan
Ushio, Asahi
Zheng, Weihua
Oh, Alice
Camacho-Collados, Jose
author_facet Ousidhoum, Nedjma
Myung, Junho
Perez-Almendros, Carla
Jin, Jiho
Keleg, Amr
Beloucif, Meriem
Zhou, Yi
Agerri, Rodrigo
Araujo, Vladimir
Baes, Naomi
Barry, James
Boisson, Joanne
Chen, Nancy F.
de Kock, Christine
Edwards, Aleksandra
de Landa, Joseba Fernandez
Imam, Mohamed Fazli
Hakami, Huda
Hsieh, Shu-Kai
Imperial, Joseph Marvin
Lee, Roy Ka-Wei
Liu, Zhengyuan
Lyu, Chenyang
Samih, Younes
Sjons, Johan
Tan, Bryan
Ushio, Asahi
Zheng, Weihua
Oh, Alice
Camacho-Collados, Jose
contents We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an extended version of our manually constructed BLEnD benchmark (Myung et al. 2024), covering more than 30 language-culture pairs, predominantly representing low-resource languages spoken across multiple continents. As the task is designed strictly for evaluation, participants were not permitted to use the data for training, fine-tuning, few-shot learning, or any other form of model modification. Our task includes two tracks: (a) Short-Answer Questions (SAQ) and (b) Multiple-Choice Questions (MCQ). Participants were required to predict labels and were allowed to submit any NLP system and adopt diverse modelling strategies, provided that the benchmark was used solely for evaluation. The task attracted more than 140 registered participants, and we received final submissions from 62 teams, along with 19 system description papers. We report the results and present an analysis of the best-performing systems and the most commonly adopted approaches. Furthermore, we discuss shared insights into open questions and challenges related to evaluation, misalignment, and methodological perspectives on model behaviour in low-resource languages and for under-represented cultures.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02601
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures
Ousidhoum, Nedjma
Myung, Junho
Perez-Almendros, Carla
Jin, Jiho
Keleg, Amr
Beloucif, Meriem
Zhou, Yi
Agerri, Rodrigo
Araujo, Vladimir
Baes, Naomi
Barry, James
Boisson, Joanne
Chen, Nancy F.
de Kock, Christine
Edwards, Aleksandra
de Landa, Joseba Fernandez
Imam, Mohamed Fazli
Hakami, Huda
Hsieh, Shu-Kai
Imperial, Joseph Marvin
Lee, Roy Ka-Wei
Liu, Zhengyuan
Lyu, Chenyang
Samih, Younes
Sjons, Johan
Tan, Bryan
Ushio, Asahi
Zheng, Weihua
Oh, Alice
Camacho-Collados, Jose
Computation and Language
We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an extended version of our manually constructed BLEnD benchmark (Myung et al. 2024), covering more than 30 language-culture pairs, predominantly representing low-resource languages spoken across multiple continents. As the task is designed strictly for evaluation, participants were not permitted to use the data for training, fine-tuning, few-shot learning, or any other form of model modification. Our task includes two tracks: (a) Short-Answer Questions (SAQ) and (b) Multiple-Choice Questions (MCQ). Participants were required to predict labels and were allowed to submit any NLP system and adopt diverse modelling strategies, provided that the benchmark was used solely for evaluation. The task attracted more than 140 registered participants, and we received final submissions from 62 teams, along with 19 system description papers. We report the results and present an analysis of the best-performing systems and the most commonly adopted approaches. Furthermore, we discuss shared insights into open questions and challenges related to evaluation, misalignment, and methodological perspectives on model behaviour in low-resource languages and for under-represented cultures.
title SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures
topic Computation and Language
url https://arxiv.org/abs/2605.02601