SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914528078331904 |
|---|---|
| author | Ousidhoum, Nedjma Myung, Junho Perez-Almendros, Carla Jin, Jiho Keleg, Amr Beloucif, Meriem Zhou, Yi Agerri, Rodrigo Araujo, Vladimir Baes, Naomi Barry, James Boisson, Joanne Chen, Nancy F. de Kock, Christine Edwards, Aleksandra de Landa, Joseba Fernandez Imam, Mohamed Fazli Hakami, Huda Hsieh, Shu-Kai Imperial, Joseph Marvin Lee, Roy Ka-Wei Liu, Zhengyuan Lyu, Chenyang Samih, Younes Sjons, Johan Tan, Bryan Ushio, Asahi Zheng, Weihua Oh, Alice Camacho-Collados, Jose |
| author_facet | Ousidhoum, Nedjma Myung, Junho Perez-Almendros, Carla Jin, Jiho Keleg, Amr Beloucif, Meriem Zhou, Yi Agerri, Rodrigo Araujo, Vladimir Baes, Naomi Barry, James Boisson, Joanne Chen, Nancy F. de Kock, Christine Edwards, Aleksandra de Landa, Joseba Fernandez Imam, Mohamed Fazli Hakami, Huda Hsieh, Shu-Kai Imperial, Joseph Marvin Lee, Roy Ka-Wei Liu, Zhengyuan Lyu, Chenyang Samih, Younes Sjons, Johan Tan, Bryan Ushio, Asahi Zheng, Weihua Oh, Alice Camacho-Collados, Jose |
| contents | We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an extended version of our manually constructed BLEnD benchmark (Myung et al. 2024), covering more than 30 language-culture pairs, predominantly representing low-resource languages spoken across multiple continents. As the task is designed strictly for evaluation, participants were not permitted to use the data for training, fine-tuning, few-shot learning, or any other form of model modification. Our task includes two tracks: (a) Short-Answer Questions (SAQ) and (b) Multiple-Choice Questions (MCQ). Participants were required to predict labels and were allowed to submit any NLP system and adopt diverse modelling strategies, provided that the benchmark was used solely for evaluation. The task attracted more than 140 registered participants, and we received final submissions from 62 teams, along with 19 system description papers. We report the results and present an analysis of the best-performing systems and the most commonly adopted approaches. Furthermore, we discuss shared insights into open questions and challenges related to evaluation, misalignment, and methodological perspectives on model behaviour in low-resource languages and for under-represented cultures. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_02601 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures Ousidhoum, Nedjma Myung, Junho Perez-Almendros, Carla Jin, Jiho Keleg, Amr Beloucif, Meriem Zhou, Yi Agerri, Rodrigo Araujo, Vladimir Baes, Naomi Barry, James Boisson, Joanne Chen, Nancy F. de Kock, Christine Edwards, Aleksandra de Landa, Joseba Fernandez Imam, Mohamed Fazli Hakami, Huda Hsieh, Shu-Kai Imperial, Joseph Marvin Lee, Roy Ka-Wei Liu, Zhengyuan Lyu, Chenyang Samih, Younes Sjons, Johan Tan, Bryan Ushio, Asahi Zheng, Weihua Oh, Alice Camacho-Collados, Jose Computation and Language We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an extended version of our manually constructed BLEnD benchmark (Myung et al. 2024), covering more than 30 language-culture pairs, predominantly representing low-resource languages spoken across multiple continents. As the task is designed strictly for evaluation, participants were not permitted to use the data for training, fine-tuning, few-shot learning, or any other form of model modification. Our task includes two tracks: (a) Short-Answer Questions (SAQ) and (b) Multiple-Choice Questions (MCQ). Participants were required to predict labels and were allowed to submit any NLP system and adopt diverse modelling strategies, provided that the benchmark was used solely for evaluation. The task attracted more than 140 registered participants, and we received final submissions from 62 teams, along with 19 system description papers. We report the results and present an analysis of the best-performing systems and the most commonly adopted approaches. Furthermore, we discuss shared insights into open questions and challenges related to evaluation, misalignment, and methodological perspectives on model behaviour in low-resource languages and for under-represented cultures. |
| title | SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2605.02601 |