An Empirical Study on the Characteristics of Bias upon Context Length Variation for Bangla

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sadhu, Jayanta, Khan, Ayan Antik, Bhattacharjee, Abhik, Shahriyar, Rifat
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929399400497152
author Sadhu, Jayanta
Khan, Ayan Antik
Bhattacharjee, Abhik
Shahriyar, Rifat
author_facet Sadhu, Jayanta
Khan, Ayan Antik
Bhattacharjee, Abhik
Shahriyar, Rifat
contents Pretrained language models inherently exhibit various social biases, prompting a crucial examination of their social impact across various linguistic contexts due to their widespread usage. Previous studies have provided numerous methods for intrinsic bias measurements, predominantly focused on high-resource languages. In this work, we aim to extend these investigations to Bangla, a low-resource language. Specifically, in this study, we (1) create a dataset for intrinsic gender bias measurement in Bangla, (2) discuss necessary adaptations to apply existing bias measurement methods for Bangla, and (3) examine the impact of context length variation on bias measurement, a factor that has been overlooked in previous studies. Through our experiments, we demonstrate a clear dependency of bias metrics on context length, highlighting the need for nuanced considerations in Bangla bias analysis. We consider our work as a stepping stone for bias measurement in the Bangla Language and make all of our resources publicly available to support future research.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17375
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Empirical Study on the Characteristics of Bias upon Context Length Variation for Bangla
Sadhu, Jayanta
Khan, Ayan Antik
Bhattacharjee, Abhik
Shahriyar, Rifat
Computation and Language
Pretrained language models inherently exhibit various social biases, prompting a crucial examination of their social impact across various linguistic contexts due to their widespread usage. Previous studies have provided numerous methods for intrinsic bias measurements, predominantly focused on high-resource languages. In this work, we aim to extend these investigations to Bangla, a low-resource language. Specifically, in this study, we (1) create a dataset for intrinsic gender bias measurement in Bangla, (2) discuss necessary adaptations to apply existing bias measurement methods for Bangla, and (3) examine the impact of context length variation on bias measurement, a factor that has been overlooked in previous studies. Through our experiments, we demonstrate a clear dependency of bias metrics on context length, highlighting the need for nuanced considerations in Bangla bias analysis. We consider our work as a stepping stone for bias measurement in the Bangla Language and make all of our resources publicly available to support future research.
title An Empirical Study on the Characteristics of Bias upon Context Length Variation for Bangla
topic Computation and Language
url https://arxiv.org/abs/2406.17375