BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Yuan, Salhan, Suchir, Caines, Andrew, Buttery, Paula, Sun, Weiwei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914107790196736
author Gao, Yuan
Salhan, Suchir
Caines, Andrew
Buttery, Paula
Sun, Weiwei
author_facet Gao, Yuan
Salhan, Suchir
Caines, Andrew
Buttery, Paula
Sun, Weiwei
contents To bridge the gap between performance-oriented benchmarks and the evaluation of cognitively inspired models, we introduce BLiSS 1.0, a Benchmark of Learner Interlingual Syntactic Structure. Our benchmark operationalizes a new paradigm of selective tolerance, testing whether a model finds a naturalistic learner error more plausible than a matched, artificial error within the same sentence. Constructed from over 2.8 million naturalistic learner sentences, BLiSS provides 136,867 controlled triplets (corrected, learner, artificial) for this purpose. Experiments on a diverse suite of models demonstrate that selective tolerance is a distinct capability from standard grammaticality, with performance clustering strongly by training paradigm. This validates BLiSS as a robust tool for measuring how different training objectives impact a model's alignment with the systematic patterns of human language acquisition.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19419
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
Gao, Yuan
Salhan, Suchir
Caines, Andrew
Buttery, Paula
Sun, Weiwei
Computation and Language
To bridge the gap between performance-oriented benchmarks and the evaluation of cognitively inspired models, we introduce BLiSS 1.0, a Benchmark of Learner Interlingual Syntactic Structure. Our benchmark operationalizes a new paradigm of selective tolerance, testing whether a model finds a naturalistic learner error more plausible than a matched, artificial error within the same sentence. Constructed from over 2.8 million naturalistic learner sentences, BLiSS provides 136,867 controlled triplets (corrected, learner, artificial) for this purpose. Experiments on a diverse suite of models demonstrate that selective tolerance is a distinct capability from standard grammaticality, with performance clustering strongly by training paradigm. This validates BLiSS as a robust tool for measuring how different training objectives impact a model's alignment with the systematic patterns of human language acquisition.
title BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
topic Computation and Language
url https://arxiv.org/abs/2510.19419