Assessing the potential of LLM-assisted annotation for corpus-based pragmatics and discourse analysis: The case of apology

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Danni, Li, Luyang, Su, Hang, Fuoli, Matteo
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916513508753408
author Yu, Danni
Li, Luyang
Su, Hang
Fuoli, Matteo
author_facet Yu, Danni
Li, Luyang
Su, Hang
Fuoli, Matteo
contents Certain forms of linguistic annotation, like part of speech and semantic tagging, can be automated with high accuracy. However, manual annotation is still necessary for complex pragmatic and discursive features that lack a direct mapping to lexical forms. This manual process is time-consuming and error-prone, limiting the scalability of function-to-form approaches in corpus linguistics. To address this, our study explores the possibility of using large language models (LLMs) to automate pragma-discursive corpus annotation. We compare GPT-3.5 (the model behind the free-to-use version of ChatGPT), GPT-4 (the model underpinning the precise mode of Bing chatbot), and a human coder in annotating apology components in English based on the local grammar framework. We find that GPT-4 outperformed GPT-3.5, with accuracy approaching that of a human coder. These results suggest that LLMs can be successfully deployed to aid pragma-discursive corpus annotation, making the process more efficient, scalable and accessible.
format Preprint
id arxiv_https___arxiv_org_abs_2305_08339
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Assessing the potential of LLM-assisted annotation for corpus-based pragmatics and discourse analysis: The case of apology
Yu, Danni
Li, Luyang
Su, Hang
Fuoli, Matteo
Computation and Language
Artificial Intelligence
Certain forms of linguistic annotation, like part of speech and semantic tagging, can be automated with high accuracy. However, manual annotation is still necessary for complex pragmatic and discursive features that lack a direct mapping to lexical forms. This manual process is time-consuming and error-prone, limiting the scalability of function-to-form approaches in corpus linguistics. To address this, our study explores the possibility of using large language models (LLMs) to automate pragma-discursive corpus annotation. We compare GPT-3.5 (the model behind the free-to-use version of ChatGPT), GPT-4 (the model underpinning the precise mode of Bing chatbot), and a human coder in annotating apology components in English based on the local grammar framework. We find that GPT-4 outperformed GPT-3.5, with accuracy approaching that of a human coder. These results suggest that LLMs can be successfully deployed to aid pragma-discursive corpus annotation, making the process more efficient, scalable and accessible.
title Assessing the potential of LLM-assisted annotation for corpus-based pragmatics and discourse analysis: The case of apology
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2305.08339