Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lin, Yi-Cheng, Lin, Tzu-Quan, Yang, Chih-Kai, Lu, Ke-Han, Chen, Wei-Chih, Kuan, Chun-Yi, Lee, Hung-yi
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916748626755584
author Lin, Yi-Cheng
Lin, Tzu-Quan
Yang, Chih-Kai
Lu, Ke-Han
Chen, Wei-Chih
Kuan, Chun-Yi
Lee, Hung-yi
author_facet Lin, Yi-Cheng
Lin, Tzu-Quan
Yang, Chih-Kai
Lu, Ke-Han
Chen, Wei-Chih
Kuan, Chun-Yi
Lee, Hung-yi
contents Speech Integrated Large Language Models (SILLMs) combine large language models with speech perception to perform diverse tasks, such as emotion recognition to speaker verification, demonstrating universal audio understanding capability. However, these models may amplify biases present in training data, potentially leading to biased access to information for marginalized groups. This work introduces a curated spoken bias evaluation toolkit and corresponding dataset. We evaluate gender bias in SILLMs across four semantic-related tasks: speech-to-text translation (STT), spoken coreference resolution (SCR), spoken sentence continuation (SSC), and spoken question answering (SQA). Our analysis reveals that bias levels are language-dependent and vary with different evaluation methods. Our findings emphasize the necessity of employing multiple approaches to comprehensively assess biases in SILLMs, providing insights for developing fairer SILLM systems.
format Preprint
id arxiv_https___arxiv_org_abs_2407_06957
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
Lin, Yi-Cheng
Lin, Tzu-Quan
Yang, Chih-Kai
Lu, Ke-Han
Chen, Wei-Chih
Kuan, Chun-Yi
Lee, Hung-yi
Audio and Speech Processing
Computation and Language
Computers and Society
Speech Integrated Large Language Models (SILLMs) combine large language models with speech perception to perform diverse tasks, such as emotion recognition to speaker verification, demonstrating universal audio understanding capability. However, these models may amplify biases present in training data, potentially leading to biased access to information for marginalized groups. This work introduces a curated spoken bias evaluation toolkit and corresponding dataset. We evaluate gender bias in SILLMs across four semantic-related tasks: speech-to-text translation (STT), spoken coreference resolution (SCR), spoken sentence continuation (SSC), and spoken question answering (SQA). Our analysis reveals that bias levels are language-dependent and vary with different evaluation methods. Our findings emphasize the necessity of employing multiple approaches to comprehensively assess biases in SILLMs, providing insights for developing fairer SILLM systems.
title Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
topic Audio and Speech Processing
Computation and Language
Computers and Society
url https://arxiv.org/abs/2407.06957