Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2507.12064 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912485918900224 |
|---|---|
| author | Ochab, Jeremi K. Matias, Mateusz Boba, Tymoteusz Walkowiak, Tomasz |
| author_facet | Ochab, Jeremi K. Matias, Mateusz Boba, Tymoteusz Walkowiak, Tomasz |
| contents | This submission to the binary AI detection task is based on a modular stylometric pipeline, where: public spaCy models are used for text preprocessing (including tokenisation, named entity recognition, dependency parsing, part-of-speech tagging, and morphology annotation) and extracting several thousand features (frequencies of n-grams of the above linguistic annotations); light-gradient boosting machines are used as the classifier. We collect a large corpus of more than 500 000 machine-generated texts for the classifier's training. We explore several parameter options to increase the classifier's capacity and take advantage of that training set. Our approach follows the non-neural, computationally inexpensive but explainable approach found effective previously. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_12064 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features Ochab, Jeremi K. Matias, Mateusz Boba, Tymoteusz Walkowiak, Tomasz Computation and Language Artificial Intelligence Machine Learning This submission to the binary AI detection task is based on a modular stylometric pipeline, where: public spaCy models are used for text preprocessing (including tokenisation, named entity recognition, dependency parsing, part-of-speech tagging, and morphology annotation) and extracting several thousand features (frequencies of n-grams of the above linguistic annotations); light-gradient boosting machines are used as the classifier. We collect a large corpus of more than 500 000 machine-generated texts for the classifier's training. We explore several parameter options to increase the classifier's capacity and take advantage of that training set. Our approach follows the non-neural, computationally inexpensive but explainable approach found effective previously. |
| title | StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2507.12064 |