TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nadas, Mihai, Diosan, Laura, Piscoran, Andrei, Tomescu, Andreea
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915972870307840
author Nadas, Mihai
Diosan, Laura
Piscoran, Andrei
Tomescu, Andreea
author_facet Nadas, Mihai
Diosan, Laura
Piscoran, Andrei
Tomescu, Andreea
contents Moral stories are a time-tested vehicle for transmitting values, yet modern NLP lacks a large, structured corpus that couples coherent narratives with explicit ethical lessons. We present TF1-EN-3M, to our knowledge the first open dataset of three million English-language fables generated exclusively by instruction-tuned models no larger than 8B parameters. Each story follows a six-slot scaffold (character -> trait -> setting -> conflict -> resolution -> moral), produced through a combinatorial prompt engine that guarantees genre fidelity while covering a broad thematic space. A fully reproducible evaluation pipeline employs a panel of open-weight LLM judges from distinct model families, scoring grammar, creativity, moral clarity, and template adherence, complemented by reference-free diversity and readability metrics. Among ten open-weight generator candidates, an 8B-parameter Llama-3 variant delivers the best quality-cost trade-off, producing high-scoring fables on consumer hardware at approximately $0.135 per 1,000 fables. We release the dataset, generation code, evaluation scripts, and full metadata under a permissive license, enabling exact reproducibility and cost benchmarking. TF1-EN-3M opens avenues for research in instruction following, narrative intelligence, value alignment, and child-friendly educational AI -- demonstrating that large-scale moral storytelling requires neither proprietary giant models nor proprietary evaluation infrastructure.
format Preprint
id arxiv_https___arxiv_org_abs_2504_20605
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
Nadas, Mihai
Diosan, Laura
Piscoran, Andrei
Tomescu, Andreea
Computation and Language
Artificial Intelligence
Digital Libraries
Machine Learning
Moral stories are a time-tested vehicle for transmitting values, yet modern NLP lacks a large, structured corpus that couples coherent narratives with explicit ethical lessons. We present TF1-EN-3M, to our knowledge the first open dataset of three million English-language fables generated exclusively by instruction-tuned models no larger than 8B parameters. Each story follows a six-slot scaffold (character -> trait -> setting -> conflict -> resolution -> moral), produced through a combinatorial prompt engine that guarantees genre fidelity while covering a broad thematic space. A fully reproducible evaluation pipeline employs a panel of open-weight LLM judges from distinct model families, scoring grammar, creativity, moral clarity, and template adherence, complemented by reference-free diversity and readability metrics. Among ten open-weight generator candidates, an 8B-parameter Llama-3 variant delivers the best quality-cost trade-off, producing high-scoring fables on consumer hardware at approximately $0.135 per 1,000 fables. We release the dataset, generation code, evaluation scripts, and full metadata under a permissive license, enabling exact reproducibility and cost benchmarking. TF1-EN-3M opens avenues for research in instruction following, narrative intelligence, value alignment, and child-friendly educational AI -- demonstrating that large-scale moral storytelling requires neither proprietary giant models nor proprietary evaluation infrastructure.
title TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
topic Computation and Language
Artificial Intelligence
Digital Libraries
Machine Learning
url https://arxiv.org/abs/2504.20605