Overview of the Plagiarism Detection Task at PAN 2025

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Greiner-Petter, André, Fröbe, Maik, Wahle, Jan Philip, Ruas, Terry, Gipp, Bela, Aizawa, Akiko, Potthast, Martin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914080451723264
author Greiner-Petter, André
Fröbe, Maik
Wahle, Jan Philip
Ruas, Terry
Gipp, Bela
Aizawa, Akiko
Potthast, Martin
author_facet Greiner-Petter, André
Fröbe, Maik
Wahle, Jan Philip
Ruas, Terry
Gipp, Bela
Aizawa, Akiko
Potthast, Martin
contents The generative plagiarism detection task at PAN 2025 aims at identifying automatically generated textual plagiarism in scientific articles and aligning them with their respective sources. We created a novel large-scale dataset of automatically generated plagiarism using three large language models: Llama, DeepSeek-R1, and Mistral. In this task overview paper, we outline the creation of this dataset, summarize and compare the results of all participants and four baselines, and evaluate the results on the last plagiarism detection task from PAN 2015 in order to interpret the robustness of the proposed approaches. We found that the current iteration does not invite a large variety of approaches as naive semantic similarity approaches based on embedding vectors provide promising results of up to 0.8 recall and 0.5 precision. In contrast, most of these approaches underperform significantly on the 2015 dataset, indicating a lack in generalizability.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Overview of the Plagiarism Detection Task at PAN 2025
Greiner-Petter, André
Fröbe, Maik
Wahle, Jan Philip
Ruas, Terry
Gipp, Bela
Aizawa, Akiko
Potthast, Martin
Computation and Language
Information Retrieval
The generative plagiarism detection task at PAN 2025 aims at identifying automatically generated textual plagiarism in scientific articles and aligning them with their respective sources. We created a novel large-scale dataset of automatically generated plagiarism using three large language models: Llama, DeepSeek-R1, and Mistral. In this task overview paper, we outline the creation of this dataset, summarize and compare the results of all participants and four baselines, and evaluate the results on the last plagiarism detection task from PAN 2015 in order to interpret the robustness of the proposed approaches. We found that the current iteration does not invite a large variety of approaches as naive semantic similarity approaches based on embedding vectors provide promising results of up to 0.8 recall and 0.5 precision. In contrast, most of these approaches underperform significantly on the 2015 dataset, indicating a lack in generalizability.
title Overview of the Plagiarism Detection Task at PAN 2025
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2510.06805