Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Utami, Nabelanita, Sasano, Ryohei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908992954957824
author Utami, Nabelanita
Sasano, Ryohei
author_facet Utami, Nabelanita
Sasano, Ryohei
contents The evolution of writing assistance tools from machine translation to large language models (LLMs) has changed how researchers write. This study investigates whether this shift is homogenizing research papers by analyzing native language identification (NLI) trends in ACL Anthology papers across three eras: pre-neural network (NN), pre-LLM, and post-LLM. We construct a labeled dataset using a semi-automated framework and fine-tune a classifier to detect linguistic fingerprints of author backgrounds. Our analysis shows a consistent decline in NLI performance over time. Interestingly, the post-LLM era reveals anomalies: while Chinese and French show unexpected resistance or divergent trends, Japanese and Korean exhibit sharper-than-expected declines.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08568
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
Utami, Nabelanita
Sasano, Ryohei
Computation and Language
Artificial Intelligence
The evolution of writing assistance tools from machine translation to large language models (LLMs) has changed how researchers write. This study investigates whether this shift is homogenizing research papers by analyzing native language identification (NLI) trends in ACL Anthology papers across three eras: pre-neural network (NN), pre-LLM, and post-LLM. We construct a labeled dataset using a semi-automated framework and fine-tune a classifier to detect linguistic fingerprints of author backgrounds. Our analysis shows a consistent decline in NLI performance over time. Interestingly, the post-LLM era reveals anomalies: while Chinese and French show unexpected resistance or divergent trends, Japanese and Korean exhibit sharper-than-expected declines.
title Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.08568