What makes a word hard to learn? Modeling L1 influence on English vocabulary difficulty

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Martins, Jonas Mayer, Huang, Zhuojing, Herygers, Aaricia, Beinborn, Lisa
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909037715521536
author Martins, Jonas Mayer
Huang, Zhuojing
Herygers, Aaricia
Beinborn, Lisa
author_facet Martins, Jonas Mayer
Huang, Zhuojing
Herygers, Aaricia
Beinborn, Lisa
contents What makes a word difficult to learn, and how does the difficulty depend on the learner's native language? We computationally model vocabulary difficulty for English learners whose first language is Spanish, German, or Chinese with gradient-boosted models trained on features related to a word's familiarity (e.g., frequency), meaning, surface form, and cross-linguistic transfer. Using Shapley values, we determine the importance of each feature group. Word familiarity is the dominant feature group shared by all three languages. However, predictions for Spanish- and German-speaking learners rely additionally on orthographic transfer. This transfer mechanism is unavailable to Chinese learners, whose difficulty is shaped by a combination of familiarity and surface features alone. Our models provide interpretable, L1-tailored difficulty estimates that can be used to design vocabulary curricula.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12281
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle What makes a word hard to learn? Modeling L1 influence on English vocabulary difficulty
Martins, Jonas Mayer
Huang, Zhuojing
Herygers, Aaricia
Beinborn, Lisa
Computation and Language
Machine Learning
What makes a word difficult to learn, and how does the difficulty depend on the learner's native language? We computationally model vocabulary difficulty for English learners whose first language is Spanish, German, or Chinese with gradient-boosted models trained on features related to a word's familiarity (e.g., frequency), meaning, surface form, and cross-linguistic transfer. Using Shapley values, we determine the importance of each feature group. Word familiarity is the dominant feature group shared by all three languages. However, predictions for Spanish- and German-speaking learners rely additionally on orthographic transfer. This transfer mechanism is unavailable to Chinese learners, whose difficulty is shaped by a combination of familiarity and surface features alone. Our models provide interpretable, L1-tailored difficulty estimates that can be used to design vocabulary curricula.
title What makes a word hard to learn? Modeling L1 influence on English vocabulary difficulty
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2605.12281