On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wan, Herun, Luo, Minnan, Su, Zhixiong, Dai, Guang, Zhao, Xiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918037457731584
author Wan, Herun
Luo, Minnan
Su, Zhixiong
Dai, Guang
Zhao, Xiang
author_facet Wan, Herun
Luo, Minnan
Su, Zhixiong
Dai, Guang
Zhao, Xiang
contents Evidence-enhanced detectors present remarkable abilities in identifying malicious social text. However, the rise of large language models (LLMs) brings potential risks of evidence pollution to confuse detectors. This paper explores potential manipulation scenarios including basic pollution, and rephrasing or generating evidence by LLMs. To mitigate the negative impact, we propose three defense strategies from the data and model sides, including machine-generated text detection, a mixture of experts, and parameter updating. Extensive experiments on four malicious social text detection tasks with ten datasets illustrate that evidence pollution significantly compromises detectors, where the generating strategy causes up to a 14.4% performance drop. Meanwhile, the defense strategies could mitigate evidence pollution, but they faced limitations for practical employment. Further analysis illustrates that polluted evidence (i) is of high quality, evaluated by metrics and humans; (ii) would compromise the model calibration, increasing expected calibration error up to 21.6%; and (iii) could be integrated to amplify the negative impact, especially for encoder-based LMs, where the accuracy drops by 21.8%.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12600
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMs
Wan, Herun
Luo, Minnan
Su, Zhixiong
Dai, Guang
Zhao, Xiang
Computation and Language
Evidence-enhanced detectors present remarkable abilities in identifying malicious social text. However, the rise of large language models (LLMs) brings potential risks of evidence pollution to confuse detectors. This paper explores potential manipulation scenarios including basic pollution, and rephrasing or generating evidence by LLMs. To mitigate the negative impact, we propose three defense strategies from the data and model sides, including machine-generated text detection, a mixture of experts, and parameter updating. Extensive experiments on four malicious social text detection tasks with ten datasets illustrate that evidence pollution significantly compromises detectors, where the generating strategy causes up to a 14.4% performance drop. Meanwhile, the defense strategies could mitigate evidence pollution, but they faced limitations for practical employment. Further analysis illustrates that polluted evidence (i) is of high quality, evaluated by metrics and humans; (ii) would compromise the model calibration, increasing expected calibration error up to 21.6%; and (iii) could be integrated to amplify the negative impact, especially for encoder-based LMs, where the accuracy drops by 21.8%.
title On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMs
topic Computation and Language
url https://arxiv.org/abs/2410.12600