Taxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nandwana, Mahesh Kumar, Lim, Youngwan, Liu, Joseph, Yang, Alex, Notibala, Varun, Khanna, Nishchaie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918233541443584
author Nandwana, Mahesh Kumar
Lim, Youngwan
Liu, Joseph
Yang, Alex
Notibala, Varun
Khanna, Nishchaie
author_facet Nandwana, Mahesh Kumar
Lim, Youngwan
Liu, Joseph
Yang, Alex
Notibala, Varun
Khanna, Nishchaie
contents Large Language Models (LLMs) are typically aligned for safety during the post-training phase; however, they may still generate inappropriate outputs that could potentially pose risks to users. This challenge underscores the need for robust safeguards that operate across both model inputs and outputs. In this work, we introduce Roblox Guard 1.0, a state-of-the-art instruction fine-tuned LLM designed to enhance the safety of LLM systems through comprehensive input-output moderation, using a pipeline of LLMs to enhance moderation capability. Built on the Llama-3.1-8B-Instruct backbone, our model is instruction fine-tuned to generalize across previously unseen safety taxonomies and demonstrates strong performance on out-of-domain safety benchmarks. The instruction fine-tuning process uses a mix of synthetic and open-source safety datasets, augmented with chain-of-thought (CoT) rationales and input inversion to enhance contextual understanding and decision making. To support systematic evaluation, we also release RobloxGuard-Eval, a new benchmark featuring an extensible safety taxonomy to assess the effectiveness of LLM guardrails and moderation frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05339
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Models
Nandwana, Mahesh Kumar
Lim, Youngwan
Liu, Joseph
Yang, Alex
Notibala, Varun
Khanna, Nishchaie
Machine Learning
Large Language Models (LLMs) are typically aligned for safety during the post-training phase; however, they may still generate inappropriate outputs that could potentially pose risks to users. This challenge underscores the need for robust safeguards that operate across both model inputs and outputs. In this work, we introduce Roblox Guard 1.0, a state-of-the-art instruction fine-tuned LLM designed to enhance the safety of LLM systems through comprehensive input-output moderation, using a pipeline of LLMs to enhance moderation capability. Built on the Llama-3.1-8B-Instruct backbone, our model is instruction fine-tuned to generalize across previously unseen safety taxonomies and demonstrates strong performance on out-of-domain safety benchmarks. The instruction fine-tuning process uses a mix of synthetic and open-source safety datasets, augmented with chain-of-thought (CoT) rationales and input inversion to enhance contextual understanding and decision making. To support systematic evaluation, we also release RobloxGuard-Eval, a new benchmark featuring an extensible safety taxonomy to assess the effectiveness of LLM guardrails and moderation frameworks.
title Taxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Models
topic Machine Learning
url https://arxiv.org/abs/2512.05339