Robust AI-Generated Text Detection by Restricted Embeddings

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kuznetsov, Kristian, Tulchinskii, Eduard, Kushnareva, Laida, Magai, German, Barannikov, Serguei, Nikolenko, Sergey, Piontkovskaya, Irina
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915196870590464
author Kuznetsov, Kristian
Tulchinskii, Eduard
Kushnareva, Laida
Magai, German
Barannikov, Serguei
Nikolenko, Sergey
Piontkovskaya, Irina
author_facet Kuznetsov, Kristian
Tulchinskii, Eduard
Kushnareva, Laida
Magai, German
Barannikov, Serguei
Nikolenko, Sergey
Piontkovskaya, Irina
contents Growing amount and quality of AI-generated texts makes detecting such content more difficult. In most real-world scenarios, the domain (style and topic) of generated data and the generator model are not known in advance. In this work, we focus on the robustness of classifier-based detectors of AI-generated text, namely their ability to transfer to unseen generators or semantic domains. We investigate the geometry of the embedding space of Transformer-based text encoders and show that clearing out harmful linear subspaces helps to train a robust classifier, ignoring domain-specific spurious features. We investigate several subspace decomposition and feature selection strategies and achieve significant improvements over state of the art methods in cross-domain and cross-generator transfer. Our best approaches for head-wise and coordinate-based subspace removal increase the mean out-of-distribution (OOD) classification score by up to 9% and 14% in particular setups for RoBERTa and BERT embeddings respectively. We release our code and data: https://github.com/SilverSolver/RobustATD
format Preprint
id arxiv_https___arxiv_org_abs_2410_08113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Robust AI-Generated Text Detection by Restricted Embeddings
Kuznetsov, Kristian
Tulchinskii, Eduard
Kushnareva, Laida
Magai, German
Barannikov, Serguei
Nikolenko, Sergey
Piontkovskaya, Irina
Computation and Language
Artificial Intelligence
Information Theory
Growing amount and quality of AI-generated texts makes detecting such content more difficult. In most real-world scenarios, the domain (style and topic) of generated data and the generator model are not known in advance. In this work, we focus on the robustness of classifier-based detectors of AI-generated text, namely their ability to transfer to unseen generators or semantic domains. We investigate the geometry of the embedding space of Transformer-based text encoders and show that clearing out harmful linear subspaces helps to train a robust classifier, ignoring domain-specific spurious features. We investigate several subspace decomposition and feature selection strategies and achieve significant improvements over state of the art methods in cross-domain and cross-generator transfer. Our best approaches for head-wise and coordinate-based subspace removal increase the mean out-of-distribution (OOD) classification score by up to 9% and 14% in particular setups for RoBERTa and BERT embeddings respectively. We release our code and data: https://github.com/SilverSolver/RobustATD
title Robust AI-Generated Text Detection by Restricted Embeddings
topic Computation and Language
Artificial Intelligence
Information Theory
url https://arxiv.org/abs/2410.08113