Active Learning for Robust and Representative LLM Generation in Safety-Critical Scenarios

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hassan, Sabit, Sicilia, Anthony, Alikhani, Malihe
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917803716509696
author Hassan, Sabit
Sicilia, Anthony
Alikhani, Malihe
author_facet Hassan, Sabit
Sicilia, Anthony
Alikhani, Malihe
contents Ensuring robust safety measures across a wide range of scenarios is crucial for user-facing systems. While Large Language Models (LLMs) can generate valuable data for safety measures, they often exhibit distributional biases, focusing on common scenarios and neglecting rare but critical cases. This can undermine the effectiveness of safety protocols developed using such data. To address this, we propose a novel framework that integrates active learning with clustering to guide LLM generation, enhancing their representativeness and robustness in safety scenarios. We demonstrate the effectiveness of our approach by constructing a dataset of 5.4K potential safety violations through an iterative process involving LLM generation and an active learner model's feedback. Our results show that the proposed framework produces a more representative set of safety scenarios without requiring prior knowledge of the underlying data distribution. Additionally, data acquired through our method improves the accuracy and F1 score of both the active learner model as well models outside the scope of active learning process, highlighting its broad applicability.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11114
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Active Learning for Robust and Representative LLM Generation in Safety-Critical Scenarios
Hassan, Sabit
Sicilia, Anthony
Alikhani, Malihe
Computation and Language
Ensuring robust safety measures across a wide range of scenarios is crucial for user-facing systems. While Large Language Models (LLMs) can generate valuable data for safety measures, they often exhibit distributional biases, focusing on common scenarios and neglecting rare but critical cases. This can undermine the effectiveness of safety protocols developed using such data. To address this, we propose a novel framework that integrates active learning with clustering to guide LLM generation, enhancing their representativeness and robustness in safety scenarios. We demonstrate the effectiveness of our approach by constructing a dataset of 5.4K potential safety violations through an iterative process involving LLM generation and an active learner model's feedback. Our results show that the proposed framework produces a more representative set of safety scenarios without requiring prior knowledge of the underlying data distribution. Additionally, data acquired through our method improves the accuracy and F1 score of both the active learner model as well models outside the scope of active learning process, highlighting its broad applicability.
title Active Learning for Robust and Representative LLM Generation in Safety-Critical Scenarios
topic Computation and Language
url https://arxiv.org/abs/2410.11114