Guardado en:
Detalles Bibliográficos
Autores principales: Ochab, Jeremi K., Matias, Mateusz, Boba, Tymoteusz, Walkowiak, Tomasz
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2507.12064
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912485918900224
author Ochab, Jeremi K.
Matias, Mateusz
Boba, Tymoteusz
Walkowiak, Tomasz
author_facet Ochab, Jeremi K.
Matias, Mateusz
Boba, Tymoteusz
Walkowiak, Tomasz
contents This submission to the binary AI detection task is based on a modular stylometric pipeline, where: public spaCy models are used for text preprocessing (including tokenisation, named entity recognition, dependency parsing, part-of-speech tagging, and morphology annotation) and extracting several thousand features (frequencies of n-grams of the above linguistic annotations); light-gradient boosting machines are used as the classifier. We collect a large corpus of more than 500 000 machine-generated texts for the classifier's training. We explore several parameter options to increase the classifier's capacity and take advantage of that training set. Our approach follows the non-neural, computationally inexpensive but explainable approach found effective previously.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12064
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features
Ochab, Jeremi K.
Matias, Mateusz
Boba, Tymoteusz
Walkowiak, Tomasz
Computation and Language
Artificial Intelligence
Machine Learning
This submission to the binary AI detection task is based on a modular stylometric pipeline, where: public spaCy models are used for text preprocessing (including tokenisation, named entity recognition, dependency parsing, part-of-speech tagging, and morphology annotation) and extracting several thousand features (frequencies of n-grams of the above linguistic annotations); light-gradient boosting machines are used as the classifier. We collect a large corpus of more than 500 000 machine-generated texts for the classifier's training. We explore several parameter options to increase the classifier's capacity and take advantage of that training set. Our approach follows the non-neural, computationally inexpensive but explainable approach found effective previously.
title StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.12064