Interpretable Coreference Resolution Evaluation Using Explicit Semantics

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Gatti, Bruno, Martinelli, Giuliano, Navigli, Roberto
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917480783413248
author Gatti, Bruno
Martinelli, Giuliano
Navigli, Roberto
author_facet Gatti, Bruno
Martinelli, Giuliano
Navigli, Roberto
contents Coreference resolution is typically evaluated using aggregate statistical metrics such as CoNLL-F1, which measure structural overlap between predicted and gold clusters. While widely used, these metrics offer limited diagnostic insights, penalizing errors without revealing whether a system struggles with specific semantic categories, such as people, locations, or events, and making it difficult to interpret model capabilities or derive actionable improvements. We address this gap by introducing a semantically-enhanced evaluation framework for coreference resolution. Our approach overlays Concept and Named Entity Recognition (CNER) onto coreference outputs, assigning semantic labels to nominal mentions and propagating them to entire coreference clusters. This enables the computation of typed scores aimed at evaluating mention extraction and linking capabilities stratified by semantic class. Across our experiments on OntoNotes, LitBank, and PreCo, we show that our framework uncovers systematic weaknesses that remain obscured by aggregate metrics. Furthermore, we demonstrate that these diagnostics can be used to design targeted, low-cost data augmentation strategies, achieving measurable out-of-domain improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10627
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Interpretable Coreference Resolution Evaluation Using Explicit Semantics
Gatti, Bruno
Martinelli, Giuliano
Navigli, Roberto
Computation and Language
Artificial Intelligence
Coreference resolution is typically evaluated using aggregate statistical metrics such as CoNLL-F1, which measure structural overlap between predicted and gold clusters. While widely used, these metrics offer limited diagnostic insights, penalizing errors without revealing whether a system struggles with specific semantic categories, such as people, locations, or events, and making it difficult to interpret model capabilities or derive actionable improvements. We address this gap by introducing a semantically-enhanced evaluation framework for coreference resolution. Our approach overlays Concept and Named Entity Recognition (CNER) onto coreference outputs, assigning semantic labels to nominal mentions and propagating them to entire coreference clusters. This enables the computation of typed scores aimed at evaluating mention extraction and linking capabilities stratified by semantic class. Across our experiments on OntoNotes, LitBank, and PreCo, we show that our framework uncovers systematic weaknesses that remain obscured by aggregate metrics. Furthermore, we demonstrate that these diagnostics can be used to design targeted, low-cost data augmentation strategies, achieving measurable out-of-domain improvements.
title Interpretable Coreference Resolution Evaluation Using Explicit Semantics
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.10627