MinorBench: A hand-built benchmark for content-based risks for children

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Khoo, Shaun, Chua, Gabriel, Shong, Rachel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909536681459712
author Khoo, Shaun
Chua, Gabriel
Shong, Rachel
author_facet Khoo, Shaun
Chua, Gabriel
Shong, Rachel
contents Large Language Models (LLMs) are rapidly entering children's lives - through parent-driven adoption, schools, and peer networks - yet current AI ethics and safety research do not adequately address content-related risks specific to minors. In this paper, we highlight these gaps with a real-world case study of an LLM-based chatbot deployed in a middle school setting, revealing how students used and sometimes misused the system. Building on these findings, we propose a new taxonomy of content-based risks for minors and introduce MinorBench, an open-source benchmark designed to evaluate LLMs on their ability to refuse unsafe or inappropriate queries from children. We evaluate six prominent LLMs under different system prompts, demonstrating substantial variability in their child-safety compliance. Our results inform practical steps for more robust, child-focused safety mechanisms and underscore the urgency of tailoring AI systems to safeguard young users.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10242
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MinorBench: A hand-built benchmark for content-based risks for children
Khoo, Shaun
Chua, Gabriel
Shong, Rachel
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) are rapidly entering children's lives - through parent-driven adoption, schools, and peer networks - yet current AI ethics and safety research do not adequately address content-related risks specific to minors. In this paper, we highlight these gaps with a real-world case study of an LLM-based chatbot deployed in a middle school setting, revealing how students used and sometimes misused the system. Building on these findings, we propose a new taxonomy of content-based risks for minors and introduce MinorBench, an open-source benchmark designed to evaluate LLMs on their ability to refuse unsafe or inappropriate queries from children. We evaluate six prominent LLMs under different system prompts, demonstrating substantial variability in their child-safety compliance. Our results inform practical steps for more robust, child-focused safety mechanisms and underscore the urgency of tailoring AI systems to safeguard young users.
title MinorBench: A hand-built benchmark for content-based risks for children
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.10242