Emergent Semantic Role Understanding in Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Griffiths, Carla, Musolesi, Mirco
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914548464746496
author Griffiths, Carla
Musolesi, Mirco
author_facet Griffiths, Carla
Musolesi, Mirco
contents Understanding how linguistic structure emerges in language models is central to interpreting what these systems learn from data and how much supervision they truly require. In particular, semantic role understanding ("who did what to whom") is a core component of meaning representation, yet it remains unclear whether it arises from pre-training alone or depends on task-specific fine-tuning. We study whether semantic role understanding emerges during language model pre-training or requires task-specific fine-tuning. We freeze decoder-only transformers and train linear probes to extract semantic roles, using performance to infer whether role information is already encoded in pre-training or learned during adaptation. Across model scales, we find that frozen representations contain substantial semantic role information, with performance improving but not fully matching fine-tuned models. This indicates partial but incomplete emergence from pre-training alone. We show that semantic role structure emerges from language modeling objectives, but its internal implementation shifts toward more distributed representations as model scale increases.
format Preprint
id arxiv_https___arxiv_org_abs_2605_09187
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Emergent Semantic Role Understanding in Language Models
Griffiths, Carla
Musolesi, Mirco
Artificial Intelligence
Computation and Language
Machine Learning
Understanding how linguistic structure emerges in language models is central to interpreting what these systems learn from data and how much supervision they truly require. In particular, semantic role understanding ("who did what to whom") is a core component of meaning representation, yet it remains unclear whether it arises from pre-training alone or depends on task-specific fine-tuning. We study whether semantic role understanding emerges during language model pre-training or requires task-specific fine-tuning. We freeze decoder-only transformers and train linear probes to extract semantic roles, using performance to infer whether role information is already encoded in pre-training or learned during adaptation. Across model scales, we find that frozen representations contain substantial semantic role information, with performance improving but not fully matching fine-tuned models. This indicates partial but incomplete emergence from pre-training alone. We show that semantic role structure emerges from language modeling objectives, but its internal implementation shifts toward more distributed representations as model scale increases.
title Emergent Semantic Role Understanding in Language Models
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2605.09187