Debiasing Multilingual LLMs in Cross-lingual Latent Space

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Peng, Qiwei, Hu, Guimin, Chai, Yekun, Søgaard, Anders
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912553118990336
author Peng, Qiwei
Hu, Guimin
Chai, Yekun
Søgaard, Anders
author_facet Peng, Qiwei
Hu, Guimin
Chai, Yekun
Søgaard, Anders
contents Debiasing techniques such as SentDebias aim to reduce bias in large language models (LLMs). Previous studies have evaluated their cross-lingual transferability by directly applying these methods to LLM representations, revealing their limited effectiveness across languages. In this work, we therefore propose to perform debiasing in a joint latent space rather than directly on LLM representations. We construct a well-aligned cross-lingual latent space using an autoencoder trained on parallel TED talk scripts. Our experiments with Aya-expanse and two debiasing techniques across four languages (English, French, German, Dutch) demonstrate that a) autoencoders effectively construct a well-aligned cross-lingual latent space, and b) applying debiasing techniques in the learned cross-lingual latent space significantly improves both the overall debiasing performance and cross-lingual transferability.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17948
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Debiasing Multilingual LLMs in Cross-lingual Latent Space
Peng, Qiwei
Hu, Guimin
Chai, Yekun
Søgaard, Anders
Computation and Language
Artificial Intelligence
Machine Learning
Debiasing techniques such as SentDebias aim to reduce bias in large language models (LLMs). Previous studies have evaluated their cross-lingual transferability by directly applying these methods to LLM representations, revealing their limited effectiveness across languages. In this work, we therefore propose to perform debiasing in a joint latent space rather than directly on LLM representations. We construct a well-aligned cross-lingual latent space using an autoencoder trained on parallel TED talk scripts. Our experiments with Aya-expanse and two debiasing techniques across four languages (English, French, German, Dutch) demonstrate that a) autoencoders effectively construct a well-aligned cross-lingual latent space, and b) applying debiasing techniques in the learned cross-lingual latent space significantly improves both the overall debiasing performance and cross-lingual transferability.
title Debiasing Multilingual LLMs in Cross-lingual Latent Space
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.17948