Detecting AI-Generated Paraphrases in Bengali: A Comparative Study of Zero-Shot and Fine-Tuned Transformers

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Islam, Md. Rakibul, Samu, Most. Sharmin Sultana, Hossain, Md. Zahid, Zaman, Farhad Uz, Bhuiyan, Md. Kamrozzaman
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915694195507200
author Islam, Md. Rakibul
Samu, Most. Sharmin Sultana
Hossain, Md. Zahid
Zaman, Farhad Uz
Bhuiyan, Md. Kamrozzaman
author_facet Islam, Md. Rakibul
Samu, Most. Sharmin Sultana
Hossain, Md. Zahid
Zaman, Farhad Uz
Bhuiyan, Md. Kamrozzaman
contents Large language models (LLMs) can produce text that closely resembles human writing. This capability raises concerns about misuse, including disinformation and content manipulation. Detecting AI-generated text is essential to maintain authenticity and prevent malicious applications. Existing research has addressed detection in multiple languages, but the Bengali language remains largely unexplored. Bengali's rich vocabulary and complex structure make distinguishing human-written and AI-generated text particularly challenging. This study investigates five transformer-based models: XLMRoBERTa-Large, mDeBERTaV3-Base, BanglaBERT-Base, IndicBERT-Base and MultilingualBERT-Base. Zero-shot evaluation shows that all models perform near chance levels (around 50% accuracy) and highlight the need for task-specific fine-tuning. Fine-tuning significantly improves performance, with XLM-RoBERTa, mDeBERTa and MultilingualBERT achieving around 91% on both accuracy and F1-score. IndicBERT demonstrates comparatively weaker performance, indicating limited effectiveness in fine-tuning for this task. This work advances AI-generated text detection in Bengali and establishes a foundation for building robust systems to counter AI-generated content.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21709
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Detecting AI-Generated Paraphrases in Bengali: A Comparative Study of Zero-Shot and Fine-Tuned Transformers
Islam, Md. Rakibul
Samu, Most. Sharmin Sultana
Hossain, Md. Zahid
Zaman, Farhad Uz
Bhuiyan, Md. Kamrozzaman
Computation and Language
Artificial Intelligence
Large language models (LLMs) can produce text that closely resembles human writing. This capability raises concerns about misuse, including disinformation and content manipulation. Detecting AI-generated text is essential to maintain authenticity and prevent malicious applications. Existing research has addressed detection in multiple languages, but the Bengali language remains largely unexplored. Bengali's rich vocabulary and complex structure make distinguishing human-written and AI-generated text particularly challenging. This study investigates five transformer-based models: XLMRoBERTa-Large, mDeBERTaV3-Base, BanglaBERT-Base, IndicBERT-Base and MultilingualBERT-Base. Zero-shot evaluation shows that all models perform near chance levels (around 50% accuracy) and highlight the need for task-specific fine-tuning. Fine-tuning significantly improves performance, with XLM-RoBERTa, mDeBERTa and MultilingualBERT achieving around 91% on both accuracy and F1-score. IndicBERT demonstrates comparatively weaker performance, indicating limited effectiveness in fine-tuning for this task. This work advances AI-generated text detection in Bengali and establishes a foundation for building robust systems to counter AI-generated content.
title Detecting AI-Generated Paraphrases in Bengali: A Comparative Study of Zero-Shot and Fine-Tuned Transformers
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.21709