SIG: Speaker Identification in Literature via Prompt-Based Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Su, Zhenlin, Xu, Liyan, Xu, Jin, Li, Jiangnan, Huangfu, Mingdu
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909111306682368
author Su, Zhenlin
Xu, Liyan
Xu, Jin
Li, Jiangnan
Huangfu, Mingdu
author_facet Su, Zhenlin
Xu, Liyan
Xu, Jin
Li, Jiangnan
Huangfu, Mingdu
contents Identifying speakers of quotations in narratives is an important task in literary analysis, with challenging scenarios including the out-of-domain inference for unseen speakers, and non-explicit cases where there are no speaker mentions in surrounding context. In this work, we propose a simple and effective approach SIG, a generation-based method that verbalizes the task and quotation input based on designed prompt templates, which also enables easy integration of other auxiliary tasks that further bolster the speaker identification performance. The prediction can either come from direct generation by the model, or be determined by the highest generation probability of each speaker candidate. Based on our approach design, SIG supports out-of-domain evaluation, and achieves open-world classification paradigm that is able to accept any forms of candidate input. We perform both cross-domain evaluation and in-domain evaluation on PDNC, the largest dataset of this task, where empirical results suggest that SIG outperforms previous baselines of complicated designs, as well as the zero-shot ChatGPT, especially excelling at those hard non-explicit scenarios by up to 17% improvement. Additional experiments on another dataset WP further corroborate the efficacy of SIG.
format Preprint
id arxiv_https___arxiv_org_abs_2312_14590
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SIG: Speaker Identification in Literature via Prompt-Based Generation
Su, Zhenlin
Xu, Liyan
Xu, Jin
Li, Jiangnan
Huangfu, Mingdu
Computation and Language
Machine Learning
Identifying speakers of quotations in narratives is an important task in literary analysis, with challenging scenarios including the out-of-domain inference for unseen speakers, and non-explicit cases where there are no speaker mentions in surrounding context. In this work, we propose a simple and effective approach SIG, a generation-based method that verbalizes the task and quotation input based on designed prompt templates, which also enables easy integration of other auxiliary tasks that further bolster the speaker identification performance. The prediction can either come from direct generation by the model, or be determined by the highest generation probability of each speaker candidate. Based on our approach design, SIG supports out-of-domain evaluation, and achieves open-world classification paradigm that is able to accept any forms of candidate input. We perform both cross-domain evaluation and in-domain evaluation on PDNC, the largest dataset of this task, where empirical results suggest that SIG outperforms previous baselines of complicated designs, as well as the zero-shot ChatGPT, especially excelling at those hard non-explicit scenarios by up to 17% improvement. Additional experiments on another dataset WP further corroborate the efficacy of SIG.
title SIG: Speaker Identification in Literature via Prompt-Based Generation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2312.14590