Detection of tortured phrases in scientific literature

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Martel, Eléna, Lentschat, Martin, Labbé, Cyril
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911771299676160
author Martel, Eléna
Lentschat, Martin
Labbé, Cyril
author_facet Martel, Eléna
Lentschat, Martin
Labbé, Cyril
contents This paper presents various automatic detection methods to extract so called tortured phrases from scientific papers. These tortured phrases, e.g. flag to clamor instead of signal to noise, are the results of paraphrasing tools used to escape plagiarism detection. We built a dataset and evaluated several strategies to flag previously undocumented tortured phrases. The proposed and tested methods are based on language models and either on embeddings similarities or on predictions of masked token. We found that an approach using token prediction and that propagates the scores to the chunk level gives the best results. With a recall value of .87 and a precision value of .61, it could retrieve new tortured phrases to be submitted to domain experts for validation.
format Preprint
id arxiv_https___arxiv_org_abs_2402_03370
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Detection of tortured phrases in scientific literature
Martel, Eléna
Lentschat, Martin
Labbé, Cyril
Information Retrieval
Artificial Intelligence
Computation and Language
Digital Libraries
This paper presents various automatic detection methods to extract so called tortured phrases from scientific papers. These tortured phrases, e.g. flag to clamor instead of signal to noise, are the results of paraphrasing tools used to escape plagiarism detection. We built a dataset and evaluated several strategies to flag previously undocumented tortured phrases. The proposed and tested methods are based on language models and either on embeddings similarities or on predictions of masked token. We found that an approach using token prediction and that propagates the scores to the chunk level gives the best results. With a recall value of .87 and a precision value of .61, it could retrieve new tortured phrases to be submitted to domain experts for validation.
title Detection of tortured phrases in scientific literature
topic Information Retrieval
Artificial Intelligence
Computation and Language
Digital Libraries
url https://arxiv.org/abs/2402.03370