HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: King, Theo, Wu, Zekun, Koshiyama, Adriano, Kazim, Emre, Treleaven, Philip
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917897745465344
author King, Theo
Wu, Zekun
Koshiyama, Adriano
Kazim, Emre
Treleaven, Philip
author_facet King, Theo
Wu, Zekun
Koshiyama, Adriano
Kazim, Emre
Treleaven, Philip
contents Stereotypes are generalised assumptions about societal groups, and even state-of-the-art LLMs using in-context learning struggle to identify them accurately. Due to the subjective nature of stereotypes, where what constitutes a stereotype can vary widely depending on cultural, social, and individual perspectives, robust explainability is crucial. Explainable models ensure that these nuanced judgments can be understood and validated by human users, promoting trust and accountability. We address these challenges by introducing HEARTS (Holistic Framework for Explainable, Sustainable, and Robust Text Stereotype Detection), a framework that enhances model performance, minimises carbon footprint, and provides transparent, interpretable explanations. We establish the Expanded Multi-Grain Stereotype Dataset (EMGSD), comprising 57,201 labelled texts across six groups, including under-represented demographics like LGBTQ+ and regional stereotypes. Ablation studies confirm that BERT models fine-tuned on EMGSD outperform those trained on individual components. We then analyse a fine-tuned, carbon-efficient ALBERT-V2 model using SHAP to generate token-level importance values, ensuring alignment with human understanding, and calculate explainability confidence scores by comparing SHAP and LIME outputs...
format Preprint
id arxiv_https___arxiv_org_abs_2409_11579
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
King, Theo
Wu, Zekun
Koshiyama, Adriano
Kazim, Emre
Treleaven, Philip
Computation and Language
Stereotypes are generalised assumptions about societal groups, and even state-of-the-art LLMs using in-context learning struggle to identify them accurately. Due to the subjective nature of stereotypes, where what constitutes a stereotype can vary widely depending on cultural, social, and individual perspectives, robust explainability is crucial. Explainable models ensure that these nuanced judgments can be understood and validated by human users, promoting trust and accountability. We address these challenges by introducing HEARTS (Holistic Framework for Explainable, Sustainable, and Robust Text Stereotype Detection), a framework that enhances model performance, minimises carbon footprint, and provides transparent, interpretable explanations. We establish the Expanded Multi-Grain Stereotype Dataset (EMGSD), comprising 57,201 labelled texts across six groups, including under-represented demographics like LGBTQ+ and regional stereotypes. Ablation studies confirm that BERT models fine-tuned on EMGSD outperform those trained on individual components. We then analyse a fine-tuned, carbon-efficient ALBERT-V2 model using SHAP to generate token-level importance values, ensuring alignment with human understanding, and calculate explainability confidence scores by comparing SHAP and LIME outputs...
title HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
topic Computation and Language
url https://arxiv.org/abs/2409.11579