Semi-Supervised Learning for Large Language Models Safety and Content Moderation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dinuta, Eduard Stefan, Sirbu, Iustin, Rebedea, Traian
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909975394123776
author Dinuta, Eduard Stefan
Sirbu, Iustin
Rebedea, Traian
author_facet Dinuta, Eduard Stefan
Sirbu, Iustin
Rebedea, Traian
contents Safety for Large Language Models (LLMs) has been an ongoing research focus since their emergence and is even more relevant nowadays with the increasing capacity of those models. Currently, there are several guardrails in place for all public LLMs and multiple proposed datasets for training safety classifiers. However, training these safety classifiers relies on large quantities of labeled data, which can be problematic to acquire, prone to labeling errors, or often include synthetic data. To address these issues, we suggest a different approach: utilizing semi-supervised learning techniques, which leverage both labeled and unlabeled data, to improve the performance on the safety task. We analyze the improvements that these techniques can offer for both prompts given to Large Language Models and the responses to those requests. Moreover, since augmentation is the central part of semi-supervised algorithms, we demonstrate the importance of using task-specific augmentations, which significantly increase the performance when compared to general-purpose augmentation techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21107
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semi-Supervised Learning for Large Language Models Safety and Content Moderation
Dinuta, Eduard Stefan
Sirbu, Iustin
Rebedea, Traian
Computation and Language
Artificial Intelligence
Machine Learning
Safety for Large Language Models (LLMs) has been an ongoing research focus since their emergence and is even more relevant nowadays with the increasing capacity of those models. Currently, there are several guardrails in place for all public LLMs and multiple proposed datasets for training safety classifiers. However, training these safety classifiers relies on large quantities of labeled data, which can be problematic to acquire, prone to labeling errors, or often include synthetic data. To address these issues, we suggest a different approach: utilizing semi-supervised learning techniques, which leverage both labeled and unlabeled data, to improve the performance on the safety task. We analyze the improvements that these techniques can offer for both prompts given to Large Language Models and the responses to those requests. Moreover, since augmentation is the central part of semi-supervised algorithms, we demonstrate the importance of using task-specific augmentations, which significantly increase the performance when compared to general-purpose augmentation techniques.
title Semi-Supervised Learning for Large Language Models Safety and Content Moderation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.21107