ChildEval: When large language models meet children's personalities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Yanyan, Han, Xue, Zhao, Chunxu, Bai, Ruiqiao, Zhang, Yaxing, Hu, Qian, Mei, Lijun, Feng, Junlan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918526384603136
author Luo, Yanyan
Han, Xue
Zhao, Chunxu
Bai, Ruiqiao
Zhang, Yaxing
Hu, Qian
Mei, Lijun
Feng, Junlan
author_facet Luo, Yanyan
Han, Xue
Zhao, Chunxu
Bai, Ruiqiao
Zhang, Yaxing
Hu, Qian
Mei, Lijun
Feng, Junlan
contents While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-specific preferences is still lacking. To address this gap, we introduce ChildEval, a benchmark for evaluating LLMs' ability to infer and follow child-centered preferences in long-context conversations. ChildEval contains 29K synthesized persona profiles of children aged 3-6, providing relatively static background information. Each persona is associated with a child preference-which may align with, conflict with, or be independent of the persona-expressed either explicitly in a single sentence or implicitly through 6-10 turn dialogues. Explicit and implicit preferences are designed to reflect the same underlying preference but differ in expression, capturing dynamic aspects of preference expression rather than changes in the static persona. The benchmark spans five top-level and fourteen sub-level categories covering children's daily lives and development. We further propose fine-grained, child-centric evaluation protocols to systematically assess open-source LLMs. Experimental results demonstrate how different personalized representations affect LLM responses and suggest that finetuning on ChildEval can enhance child-centered performance. Our code and dataset are available at https://github.com/ziyanluo/ChildEval.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27805
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ChildEval: When large language models meet children's personalities
Luo, Yanyan
Han, Xue
Zhao, Chunxu
Bai, Ruiqiao
Zhang, Yaxing
Hu, Qian
Mei, Lijun
Feng, Junlan
Computation and Language
Artificial Intelligence
While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-specific preferences is still lacking. To address this gap, we introduce ChildEval, a benchmark for evaluating LLMs' ability to infer and follow child-centered preferences in long-context conversations. ChildEval contains 29K synthesized persona profiles of children aged 3-6, providing relatively static background information. Each persona is associated with a child preference-which may align with, conflict with, or be independent of the persona-expressed either explicitly in a single sentence or implicitly through 6-10 turn dialogues. Explicit and implicit preferences are designed to reflect the same underlying preference but differ in expression, capturing dynamic aspects of preference expression rather than changes in the static persona. The benchmark spans five top-level and fourteen sub-level categories covering children's daily lives and development. We further propose fine-grained, child-centric evaluation protocols to systematically assess open-source LLMs. Experimental results demonstrate how different personalized representations affect LLM responses and suggest that finetuning on ChildEval can enhance child-centered performance. Our code and dataset are available at https://github.com/ziyanluo/ChildEval.
title ChildEval: When large language models meet children's personalities
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.27805