IMGTB: A Framework for Machine-Generated Text Detection Benchmarking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Spiegel, Michal, Macko, Dominik
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912159379750912
author Spiegel, Michal
Macko, Dominik
author_facet Spiegel, Michal
Macko, Dominik
contents In the era of large language models generating high quality texts, it is a necessity to develop methods for detection of machine-generated text to avoid harmful use or simply due to annotation purposes. It is, however, also important to properly evaluate and compare such developed methods. Recently, a few benchmarks have been proposed for this purpose; however, integration of newest detection methods is rather challenging, since new methods appear each month and provide slightly different evaluation pipelines. In this paper, we present the IMGTB framework, which simplifies the benchmarking of machine-generated text detection methods by easy integration of custom (new) methods and evaluation datasets. Its configurability and flexibility makes research and development of new detection methods easier, especially their comparison to the existing state-of-the-art detectors. The default set of analyses, metrics and visualizations offered by the tool follows the established practices of machine-generated text detection benchmarking found in state-of-the-art literature.
format Preprint
id arxiv_https___arxiv_org_abs_2311_12574
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle IMGTB: A Framework for Machine-Generated Text Detection Benchmarking
Spiegel, Michal
Macko, Dominik
Computation and Language
Artificial Intelligence
In the era of large language models generating high quality texts, it is a necessity to develop methods for detection of machine-generated text to avoid harmful use or simply due to annotation purposes. It is, however, also important to properly evaluate and compare such developed methods. Recently, a few benchmarks have been proposed for this purpose; however, integration of newest detection methods is rather challenging, since new methods appear each month and provide slightly different evaluation pipelines. In this paper, we present the IMGTB framework, which simplifies the benchmarking of machine-generated text detection methods by easy integration of custom (new) methods and evaluation datasets. Its configurability and flexibility makes research and development of new detection methods easier, especially their comparison to the existing state-of-the-art detectors. The default set of analyses, metrics and visualizations offered by the tool follows the established practices of machine-generated text detection benchmarking found in state-of-the-art literature.
title IMGTB: A Framework for Machine-Generated Text Detection Benchmarking
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2311.12574