Learning Constituent Headedness

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qi, Zeyao, Chen, Yige, Lim, KyungTae, Pan, Haihua, Park, Jungyeul
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915865240272896
author Qi, Zeyao
Chen, Yige
Lim, KyungTae
Pan, Haihua
Park, Jungyeul
author_facet Qi, Zeyao
Chen, Yige
Lim, KyungTae
Pan, Haihua
Park, Jungyeul
contents Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurally via percolation rules. We treat this notion of constituent headedness as an explicit representational layer and learn it as a supervised prediction task over aligned constituency and dependency annotations, inducing supervision by defining each constituent head as the dependency span head. On aligned English and Chinese data, the resulting models achieve near-ceiling intrinsic accuracy and substantially outperform Collins-style rule-based percolation. Predicted heads yield comparable parsing accuracy under head-driven binarization, consistent with the induced binary training targets being largely equivalent across head choices, while increasing the fidelity of deterministic constituency-to-dependency conversion and transferring across resources and languages under simple label-mapping interfaces.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14755
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning Constituent Headedness
Qi, Zeyao
Chen, Yige
Lim, KyungTae
Pan, Haihua
Park, Jungyeul
Computation and Language
Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurally via percolation rules. We treat this notion of constituent headedness as an explicit representational layer and learn it as a supervised prediction task over aligned constituency and dependency annotations, inducing supervision by defining each constituent head as the dependency span head. On aligned English and Chinese data, the resulting models achieve near-ceiling intrinsic accuracy and substantially outperform Collins-style rule-based percolation. Predicted heads yield comparable parsing accuracy under head-driven binarization, consistent with the induced binary training targets being largely equivalent across head choices, while increasing the fidelity of deterministic constituency-to-dependency conversion and transferring across resources and languages under simple label-mapping interfaces.
title Learning Constituent Headedness
topic Computation and Language
url https://arxiv.org/abs/2603.14755