KidLM: Advancing Language Models for Children -- Early Insights and Future Directions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nayeem, Mir Tafseer, Rafiei, Davood
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914965360738304
author Nayeem, Mir Tafseer
Rafiei, Davood
author_facet Nayeem, Mir Tafseer
Rafiei, Davood
contents Recent studies highlight the potential of large language models in creating educational tools for children, yet significant challenges remain in maintaining key child-specific properties such as linguistic nuances, cognitive needs, and safety standards. In this paper, we explore foundational steps toward the development of child-specific language models, emphasizing the necessity of high-quality pre-training data. We introduce a novel user-centric data collection pipeline that involves gathering and validating a corpus specifically written for and sometimes by children. Additionally, we propose a new training objective, Stratified Masking, which dynamically adjusts masking probabilities based on our domain-specific child language data, enabling models to prioritize vocabulary and concepts more suitable for children. Experimental evaluations demonstrate that our model excels in understanding lower grade-level text, maintains safety by avoiding stereotypes, and captures children's unique preferences. Furthermore, we provide actionable insights for future research and development in child-specific language modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03884
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle KidLM: Advancing Language Models for Children -- Early Insights and Future Directions
Nayeem, Mir Tafseer
Rafiei, Davood
Computation and Language
Artificial Intelligence
Computers and Society
Human-Computer Interaction
Recent studies highlight the potential of large language models in creating educational tools for children, yet significant challenges remain in maintaining key child-specific properties such as linguistic nuances, cognitive needs, and safety standards. In this paper, we explore foundational steps toward the development of child-specific language models, emphasizing the necessity of high-quality pre-training data. We introduce a novel user-centric data collection pipeline that involves gathering and validating a corpus specifically written for and sometimes by children. Additionally, we propose a new training objective, Stratified Masking, which dynamically adjusts masking probabilities based on our domain-specific child language data, enabling models to prioritize vocabulary and concepts more suitable for children. Experimental evaluations demonstrate that our model excels in understanding lower grade-level text, maintains safety by avoiding stereotypes, and captures children's unique preferences. Furthermore, we provide actionable insights for future research and development in child-specific language modeling.
title KidLM: Advancing Language Models for Children -- Early Insights and Future Directions
topic Computation and Language
Artificial Intelligence
Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2410.03884