Enhancing Automatic Term Extraction with Large Language Models via Syntactic Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chun, Yongchan, Kim, Minhyuk, Kim, Dongjun, Park, Chanjun, Lim, Heuiseok
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915360018530304
author Chun, Yongchan
Kim, Minhyuk
Kim, Dongjun
Park, Chanjun
Lim, Heuiseok
author_facet Chun, Yongchan
Kim, Minhyuk
Kim, Dongjun
Park, Chanjun
Lim, Heuiseok
contents Automatic Term Extraction (ATE) identifies domain-specific expressions that are crucial for downstream tasks such as machine translation and information retrieval. Although large language models (LLMs) have significantly advanced various NLP tasks, their potential for ATE has scarcely been examined. We propose a retrieval-based prompting strategy that, in the few-shot setting, selects demonstrations according to \emph{syntactic} rather than semantic similarity. This syntactic retrieval method is domain-agnostic and provides more reliable guidance for capturing term boundaries. We evaluate the approach in both in-domain and cross-domain settings, analyzing how lexical overlap between the query sentence and its retrieved examples affects performance. Experiments on three specialized ATE benchmarks show that syntactic retrieval improves F1-score. These findings highlight the importance of syntactic cues when adapting LLMs to terminology-extraction tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Automatic Term Extraction with Large Language Models via Syntactic Retrieval
Chun, Yongchan
Kim, Minhyuk
Kim, Dongjun
Park, Chanjun
Lim, Heuiseok
Computation and Language
Information Retrieval
Automatic Term Extraction (ATE) identifies domain-specific expressions that are crucial for downstream tasks such as machine translation and information retrieval. Although large language models (LLMs) have significantly advanced various NLP tasks, their potential for ATE has scarcely been examined. We propose a retrieval-based prompting strategy that, in the few-shot setting, selects demonstrations according to \emph{syntactic} rather than semantic similarity. This syntactic retrieval method is domain-agnostic and provides more reliable guidance for capturing term boundaries. We evaluate the approach in both in-domain and cross-domain settings, analyzing how lexical overlap between the query sentence and its retrieved examples affects performance. Experiments on three specialized ATE benchmarks show that syntactic retrieval improves F1-score. These findings highlight the importance of syntactic cues when adapting LLMs to terminology-extraction tasks.
title Enhancing Automatic Term Extraction with Large Language Models via Syntactic Retrieval
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2506.21222