Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yamaguchi, Atsuki, Mi, Maggie, Aletras, Nikolaos
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915939121889280
author Yamaguchi, Atsuki
Mi, Maggie
Aletras, Nikolaos
author_facet Yamaguchi, Atsuki
Mi, Maggie
Aletras, Nikolaos
contents Language models (LMs) are pre-trained on raw text datasets to generate text sequences token-by-token. While this approach facilitates the learning of world knowledge and reasoning, it does not explicitly optimize for linguistic competence. To bridge this gap, we propose L2T, a pre-training framework integrating Language Learning Tasks alongside standard next-token prediction. Inspired by human language acquisition, L2T transforms raw text into structured input-output pairs to provide explicit linguistic stimulation. Pre-training LMs on a mixture of raw text and L2T data not only improves overall performance on linguistic competence benchmarks but accelerates its acquisition, while maintaining competitive performance on general reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03448
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
Yamaguchi, Atsuki
Mi, Maggie
Aletras, Nikolaos
Computation and Language
Language models (LMs) are pre-trained on raw text datasets to generate text sequences token-by-token. While this approach facilitates the learning of world knowledge and reasoning, it does not explicitly optimize for linguistic competence. To bridge this gap, we propose L2T, a pre-training framework integrating Language Learning Tasks alongside standard next-token prediction. Inspired by human language acquisition, L2T transforms raw text into structured input-output pairs to provide explicit linguistic stimulation. Pre-training LMs on a mixture of raw text and L2T data not only improves overall performance on linguistic competence benchmarks but accelerates its acquisition, while maintaining competitive performance on general reasoning tasks.
title Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
topic Computation and Language
url https://arxiv.org/abs/2601.03448