The power of text similarity in identifying AI-LLM paraphrased documents: The case of BBC news articles and ChatGPT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xylogiannopoulos, Konstantinos, Xanthopoulos, Petros, Karampelas, Panagiotis, Bakamitsos, Georgios
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916946911428608
author Xylogiannopoulos, Konstantinos
Xanthopoulos, Petros
Karampelas, Panagiotis
Bakamitsos, Georgios
author_facet Xylogiannopoulos, Konstantinos
Xanthopoulos, Petros
Karampelas, Panagiotis
Bakamitsos, Georgios
contents Generative AI paraphrased text can be used for copyright infringement and the AI paraphrased content can deprive substantial revenue from original content creators. Despite this recent surge of malicious use of generative AI, there are few academic publications that research this threat. In this article, we demonstrate the ability of pattern-based similarity detection for AI paraphrased news recognition. We propose an algorithmic scheme, which is not limited to detect whether an article is an AI paraphrase, but, more importantly, to identify that the source of infringement is the ChatGPT. The proposed method is tested with a benchmark dataset specifically created for this task that incorporates real articles from BBC, incorporating a total of 2,224 articles across five different news categories, as well as 2,224 paraphrased articles created with ChatGPT. Results show that our pattern similarity-based method, that makes no use of deep learning, can detect ChatGPT assisted paraphrased articles at percentages 96.23% for accuracy, 96.25% for precision, 96.21% for sensitivity, 96.25% for specificity and 96.23% for F1 score.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The power of text similarity in identifying AI-LLM paraphrased documents: The case of BBC news articles and ChatGPT
Xylogiannopoulos, Konstantinos
Xanthopoulos, Petros
Karampelas, Panagiotis
Bakamitsos, Georgios
Computation and Language
Artificial Intelligence
Generative AI paraphrased text can be used for copyright infringement and the AI paraphrased content can deprive substantial revenue from original content creators. Despite this recent surge of malicious use of generative AI, there are few academic publications that research this threat. In this article, we demonstrate the ability of pattern-based similarity detection for AI paraphrased news recognition. We propose an algorithmic scheme, which is not limited to detect whether an article is an AI paraphrase, but, more importantly, to identify that the source of infringement is the ChatGPT. The proposed method is tested with a benchmark dataset specifically created for this task that incorporates real articles from BBC, incorporating a total of 2,224 articles across five different news categories, as well as 2,224 paraphrased articles created with ChatGPT. Results show that our pattern similarity-based method, that makes no use of deep learning, can detect ChatGPT assisted paraphrased articles at percentages 96.23% for accuracy, 96.25% for precision, 96.21% for sensitivity, 96.25% for specificity and 96.23% for F1 score.
title The power of text similarity in identifying AI-LLM paraphrased documents: The case of BBC news articles and ChatGPT
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.12405