Not-in-Perspective: Towards Shielding Google's Perspective API Against Adversarial Negation Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alexiou, Michail S., Mertoguno, J. Sukarno
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908825258295296
author Alexiou, Michail S.
Mertoguno, J. Sukarno
author_facet Alexiou, Michail S.
Mertoguno, J. Sukarno
contents The rise of cyberbullying in social media platforms involving toxic comments has escalated the need for effective ways to monitor and moderate online interactions. Existing solutions of automated toxicity detection systems, are based on a machine or deep learning algorithms. However, statistics-based solutions are generally prone to adversarial attacks that contain logic based modifications such as negation in phrases and sentences. In that regard, we present a set of formal reasoning-based methodologies that wrap around existing machine learning toxicity detection systems. Acting as both pre-processing and post-processing steps, our formal reasoning wrapper helps alleviating the negation attack problems and significantly improves the accuracy and efficacy of toxicity scoring. We evaluate different variations of our wrapper on multiple machine learning models against a negation adversarial dataset. Experimental results highlight the improvement of hybrid (formal reasoning and machine-learning) methods against various purely statistical solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09343
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Not-in-Perspective: Towards Shielding Google's Perspective API Against Adversarial Negation Attacks
Alexiou, Michail S.
Mertoguno, J. Sukarno
Artificial Intelligence
Computation and Language
The rise of cyberbullying in social media platforms involving toxic comments has escalated the need for effective ways to monitor and moderate online interactions. Existing solutions of automated toxicity detection systems, are based on a machine or deep learning algorithms. However, statistics-based solutions are generally prone to adversarial attacks that contain logic based modifications such as negation in phrases and sentences. In that regard, we present a set of formal reasoning-based methodologies that wrap around existing machine learning toxicity detection systems. Acting as both pre-processing and post-processing steps, our formal reasoning wrapper helps alleviating the negation attack problems and significantly improves the accuracy and efficacy of toxicity scoring. We evaluate different variations of our wrapper on multiple machine learning models against a negation adversarial dataset. Experimental results highlight the improvement of hybrid (formal reasoning and machine-learning) methods against various purely statistical solutions.
title Not-in-Perspective: Towards Shielding Google's Perspective API Against Adversarial Negation Attacks
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.09343