Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908992954957824 |
|---|---|
| author | Utami, Nabelanita Sasano, Ryohei |
| author_facet | Utami, Nabelanita Sasano, Ryohei |
| contents | The evolution of writing assistance tools from machine translation to large language models (LLMs) has changed how researchers write. This study investigates whether this shift is homogenizing research papers by analyzing native language identification (NLI) trends in ACL Anthology papers across three eras: pre-neural network (NN), pre-LLM, and post-LLM. We construct a labeled dataset using a semi-automated framework and fine-tune a classifier to detect linguistic fingerprints of author backgrounds. Our analysis shows a consistent decline in NLI performance over time. Interestingly, the post-LLM era reveals anomalies: while Chinese and French show unexpected resistance or divergent trends, Japanese and Korean exhibit sharper-than-expected declines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_08568 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era Utami, Nabelanita Sasano, Ryohei Computation and Language Artificial Intelligence The evolution of writing assistance tools from machine translation to large language models (LLMs) has changed how researchers write. This study investigates whether this shift is homogenizing research papers by analyzing native language identification (NLI) trends in ACL Anthology papers across three eras: pre-neural network (NN), pre-LLM, and post-LLM. We construct a labeled dataset using a semi-automated framework and fine-tune a classifier to detect linguistic fingerprints of author backgrounds. Our analysis shows a consistent decline in NLI performance over time. Interestingly, the post-LLM era reveals anomalies: while Chinese and French show unexpected resistance or divergent trends, Japanese and Korean exhibit sharper-than-expected declines. |
| title | Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2604.08568 |