PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Michail, Andrianos, Clematide, Simon, Opitz, Juri
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909429542158336
author Michail, Andrianos
Clematide, Simon
Opitz, Juri
author_facet Michail, Andrianos
Clematide, Simon
Opitz, Juri
contents The task of determining whether two texts are paraphrases has long been a challenge in NLP. However, the prevailing notion of paraphrase is often quite simplistic, offering only a limited view of the vast spectrum of paraphrase phenomena. Indeed, we find that evaluating models in a paraphrase dataset can leave uncertainty about their true semantic understanding. To alleviate this, we create PARAPHRASUS, a benchmark designed for multi-dimensional assessment, benchmarking and selection of paraphrase detection models. We find that paraphrase detection models under our fine-grained evaluation lens exhibit trade-offs that cannot be captured through a single classification dataset. Furthermore, PARAPHRASUS allows prompt calibration for different use cases, tailoring LLM models to specific strictness levels. PARAPHRASUS includes 3 challenges spanning over 10 datasets, including 8 repurposed and 2 newly annotated; we release it along with a benchmarking library at https://github.com/impresso/paraphrasus
format Preprint
id arxiv_https___arxiv_org_abs_2409_12060
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
Michail, Andrianos
Clematide, Simon
Opitz, Juri
Computation and Language
Artificial Intelligence
I.2.7
The task of determining whether two texts are paraphrases has long been a challenge in NLP. However, the prevailing notion of paraphrase is often quite simplistic, offering only a limited view of the vast spectrum of paraphrase phenomena. Indeed, we find that evaluating models in a paraphrase dataset can leave uncertainty about their true semantic understanding. To alleviate this, we create PARAPHRASUS, a benchmark designed for multi-dimensional assessment, benchmarking and selection of paraphrase detection models. We find that paraphrase detection models under our fine-grained evaluation lens exhibit trade-offs that cannot be captured through a single classification dataset. Furthermore, PARAPHRASUS allows prompt calibration for different use cases, tailoring LLM models to specific strictness levels. PARAPHRASUS includes 3 challenges spanning over 10 datasets, including 8 repurposed and 2 newly annotated; we release it along with a benchmarking library at https://github.com/impresso/paraphrasus
title PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2409.12060