The Multilingual Divide and Its Impact on Global AI Safety

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peppin, Aidan, Kreutzer, Julia, Sebag, Alice Schoenauer, Marchisio, Kelly, Ermis, Beyza, Dang, John, Cahyawijaya, Samuel, Singh, Shivalika, Goldfarb-Tarrant, Seraphina, Aryabumi, Viraat, Aakanksha, Ko, Wei-Yin, Üstün, Ahmet, Gallé, Matthias, Fadaee, Marzieh, Hooker, Sara
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910970745454592
author Peppin, Aidan
Kreutzer, Julia
Sebag, Alice Schoenauer
Marchisio, Kelly
Ermis, Beyza
Dang, John
Cahyawijaya, Samuel
Singh, Shivalika
Goldfarb-Tarrant, Seraphina
Aryabumi, Viraat
Aakanksha
Ko, Wei-Yin
Üstün, Ahmet
Gallé, Matthias
Fadaee, Marzieh
Hooker, Sara
author_facet Peppin, Aidan
Kreutzer, Julia
Sebag, Alice Schoenauer
Marchisio, Kelly
Ermis, Beyza
Dang, John
Cahyawijaya, Samuel
Singh, Shivalika
Goldfarb-Tarrant, Seraphina
Aryabumi, Viraat
Aakanksha
Ko, Wei-Yin
Üstün, Ahmet
Gallé, Matthias
Fadaee, Marzieh
Hooker, Sara
contents Despite advances in large language model capabilities in recent years, a large gap remains in their capabilities and safety performance for many languages beyond a relatively small handful of globally dominant languages. This paper provides researchers, policymakers and governance experts with an overview of key challenges to bridging the "language gap" in AI and minimizing safety risks across languages. We provide an analysis of why the language gap in AI exists and grows, and how it creates disparities in global AI safety. We identify barriers to address these challenges, and recommend how those working in policy and governance can help address safety concerns associated with the language gap by supporting multilingual dataset creation, transparency, and research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21344
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Multilingual Divide and Its Impact on Global AI Safety
Peppin, Aidan
Kreutzer, Julia
Sebag, Alice Schoenauer
Marchisio, Kelly
Ermis, Beyza
Dang, John
Cahyawijaya, Samuel
Singh, Shivalika
Goldfarb-Tarrant, Seraphina
Aryabumi, Viraat
Aakanksha
Ko, Wei-Yin
Üstün, Ahmet
Gallé, Matthias
Fadaee, Marzieh
Hooker, Sara
Artificial Intelligence
Computation and Language
Despite advances in large language model capabilities in recent years, a large gap remains in their capabilities and safety performance for many languages beyond a relatively small handful of globally dominant languages. This paper provides researchers, policymakers and governance experts with an overview of key challenges to bridging the "language gap" in AI and minimizing safety risks across languages. We provide an analysis of why the language gap in AI exists and grows, and how it creates disparities in global AI safety. We identify barriers to address these challenges, and recommend how those working in policy and governance can help address safety concerns associated with the language gap by supporting multilingual dataset creation, transparency, and research.
title The Multilingual Divide and Its Impact on Global AI Safety
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.21344