SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xing, Wenpeng, Wei, Lanyi, Hu, Haixiao, Yu, Jingyi, Li, Rongchang, Li, Mohan, Lin, Changting, Han, Meng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908707835609088
author Xing, Wenpeng
Wei, Lanyi
Hu, Haixiao
Yu, Jingyi
Li, Rongchang
Li, Mohan
Lin, Changting
Han, Meng
author_facet Xing, Wenpeng
Wei, Lanyi
Hu, Haixiao
Yu, Jingyi
Li, Rongchang
Li, Mohan
Lin, Changting
Han, Meng
contents The rapid proliferation of large language models (LLMs) in applications targeting children and adolescents necessitates a fundamental reassessment of prevailing AI safety frameworks, which are largely tailored to adult users and neglect the distinct developmental vulnerabilities of minors. This paper highlights key deficiencies in existing LLM safety benchmarks, including their inadequate coverage of age-specific cognitive, emotional, and social risks spanning early childhood (ages 0--6), middle childhood (7--12), and adolescence (13--18). To bridge these gaps, we introduce SproutBench, an innovative evaluation suite comprising 1,283 developmentally grounded adversarial prompts designed to probe risks such as emotional dependency, privacy violations, and imitation of hazardous behaviors. Through rigorous empirical evaluation of 47 diverse LLMs, we uncover substantial safety vulnerabilities, corroborated by robust inter-dimensional correlations (e.g., between Safety and Risk Prevention) and a notable inverse relationship between Interactivity and Age Appropriateness. These insights yield practical guidelines for advancing child-centric AI design and deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth
Xing, Wenpeng
Wei, Lanyi
Hu, Haixiao
Yu, Jingyi
Li, Rongchang
Li, Mohan
Lin, Changting
Han, Meng
Computation and Language
Artificial Intelligence
The rapid proliferation of large language models (LLMs) in applications targeting children and adolescents necessitates a fundamental reassessment of prevailing AI safety frameworks, which are largely tailored to adult users and neglect the distinct developmental vulnerabilities of minors. This paper highlights key deficiencies in existing LLM safety benchmarks, including their inadequate coverage of age-specific cognitive, emotional, and social risks spanning early childhood (ages 0--6), middle childhood (7--12), and adolescence (13--18). To bridge these gaps, we introduce SproutBench, an innovative evaluation suite comprising 1,283 developmentally grounded adversarial prompts designed to probe risks such as emotional dependency, privacy violations, and imitation of hazardous behaviors. Through rigorous empirical evaluation of 47 diverse LLMs, we uncover substantial safety vulnerabilities, corroborated by robust inter-dimensional correlations (e.g., between Safety and Risk Prevention) and a notable inverse relationship between Interactivity and Age Appropriateness. These insights yield practical guidelines for advancing child-centric AI design and deployment.
title SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.11009