Fine-tuning Large Language Models for Multigenerator, Multidomain, and Multilingual Machine-Generated Text Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiong, Feng, Markchom, Thanet, Zheng, Ziwei, Jung, Subin, Ojha, Varun, Liang, Huizhi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911762970836992
author Xiong, Feng
Markchom, Thanet
Zheng, Ziwei
Jung, Subin
Ojha, Varun
Liang, Huizhi
author_facet Xiong, Feng
Markchom, Thanet
Zheng, Ziwei
Jung, Subin
Ojha, Varun
Liang, Huizhi
contents SemEval-2024 Task 8 introduces the challenge of identifying machine-generated texts from diverse Large Language Models (LLMs) in various languages and domains. The task comprises three subtasks: binary classification in monolingual and multilingual (Subtask A), multi-class classification (Subtask B), and mixed text detection (Subtask C). This paper focuses on Subtask A & B. Each subtask is supported by three datasets for training, development, and testing. To tackle this task, two methods: 1) using traditional machine learning (ML) with natural language preprocessing (NLP) for feature extraction, and 2) fine-tuning LLMs for text classification. The results show that transformer models, particularly LoRA-RoBERTa, exceed traditional ML methods in effectiveness, with majority voting being particularly effective in multilingual contexts for identifying machine-generated texts.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12326
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fine-tuning Large Language Models for Multigenerator, Multidomain, and Multilingual Machine-Generated Text Detection
Xiong, Feng
Markchom, Thanet
Zheng, Ziwei
Jung, Subin
Ojha, Varun
Liang, Huizhi
Computation and Language
Artificial Intelligence
SemEval-2024 Task 8 introduces the challenge of identifying machine-generated texts from diverse Large Language Models (LLMs) in various languages and domains. The task comprises three subtasks: binary classification in monolingual and multilingual (Subtask A), multi-class classification (Subtask B), and mixed text detection (Subtask C). This paper focuses on Subtask A & B. Each subtask is supported by three datasets for training, development, and testing. To tackle this task, two methods: 1) using traditional machine learning (ML) with natural language preprocessing (NLP) for feature extraction, and 2) fine-tuning LLMs for text classification. The results show that transformer models, particularly LoRA-RoBERTa, exceed traditional ML methods in effectiveness, with majority voting being particularly effective in multilingual contexts for identifying machine-generated texts.
title Fine-tuning Large Language Models for Multigenerator, Multidomain, and Multilingual Machine-Generated Text Detection
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.12326