SEFD: Semantic-Enhanced Framework for Detecting LLM-Generated Text

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: He, Weiqing, Hou, Bojian, Shang, Tianqi, Tarzanagh, Davoud Ataee, Long, Qi, Shen, Li
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909396974436352
author He, Weiqing
Hou, Bojian
Shang, Tianqi
Tarzanagh, Davoud Ataee
Long, Qi
Shen, Li
author_facet He, Weiqing
Hou, Bojian
Shang, Tianqi
Tarzanagh, Davoud Ataee
Long, Qi
Shen, Li
contents The widespread adoption of large language models (LLMs) has created an urgent need for robust tools to detect LLM-generated text, especially in light of \textit{paraphrasing} techniques that often evade existing detection methods. To address this challenge, we present a novel semantic-enhanced framework for detecting LLM-generated text (SEFD) that leverages a retrieval-based mechanism to fully utilize text semantics. Our framework improves upon existing detection methods by systematically integrating retrieval-based techniques with traditional detectors, employing a carefully curated retrieval mechanism that strikes a balance between comprehensive coverage and computational efficiency. We showcase the effectiveness of our approach in sequential text scenarios common in real-world applications, such as online forums and Q\&A platforms. Through comprehensive experiments across various LLM-generated texts and detection methods, we demonstrate that our framework substantially enhances detection accuracy in paraphrasing scenarios while maintaining robustness for standard LLM-generated content.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12764
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SEFD: Semantic-Enhanced Framework for Detecting LLM-Generated Text
He, Weiqing
Hou, Bojian
Shang, Tianqi
Tarzanagh, Davoud Ataee
Long, Qi
Shen, Li
Computation and Language
Artificial Intelligence
Information Retrieval
The widespread adoption of large language models (LLMs) has created an urgent need for robust tools to detect LLM-generated text, especially in light of \textit{paraphrasing} techniques that often evade existing detection methods. To address this challenge, we present a novel semantic-enhanced framework for detecting LLM-generated text (SEFD) that leverages a retrieval-based mechanism to fully utilize text semantics. Our framework improves upon existing detection methods by systematically integrating retrieval-based techniques with traditional detectors, employing a carefully curated retrieval mechanism that strikes a balance between comprehensive coverage and computational efficiency. We showcase the effectiveness of our approach in sequential text scenarios common in real-world applications, such as online forums and Q\&A platforms. Through comprehensive experiments across various LLM-generated texts and detection methods, we demonstrate that our framework substantially enhances detection accuracy in paraphrasing scenarios while maintaining robustness for standard LLM-generated content.
title SEFD: Semantic-Enhanced Framework for Detecting LLM-Generated Text
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2411.12764