SENTRA: Selected-Next-Token Transformer for LLM Text Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915502983479296 |
|---|---|
| author | Plyler, Mitchell Zhang, Yilun Tuzhilin, Alexander Khalifah, Saoud Tian, Sen |
| author_facet | Plyler, Mitchell Zhang, Yilun Tuzhilin, Alexander Khalifah, Saoud Tian, Sen |
| contents | LLMs are becoming increasingly capable and widespread. Consequently, the potential and reality of their misuse is also growing. In this work, we address the problem of detecting LLM-generated text that is not explicitly declared as such. We present a novel, general-purpose, and supervised LLM text detector, SElected-Next-Token tRAnsformer (SENTRA). SENTRA is a Transformer-based encoder leveraging selected-next-token-probability sequences and utilizing contrastive pre-training on large amounts of unlabeled data. Our experiments on three popular public datasets across 24 domains of text demonstrate SENTRA is a general-purpose classifier that significantly outperforms popular baselines in the out-of-domain setting. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_12385 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SENTRA: Selected-Next-Token Transformer for LLM Text Detection Plyler, Mitchell Zhang, Yilun Tuzhilin, Alexander Khalifah, Saoud Tian, Sen Computation and Language Machine Learning LLMs are becoming increasingly capable and widespread. Consequently, the potential and reality of their misuse is also growing. In this work, we address the problem of detecting LLM-generated text that is not explicitly declared as such. We present a novel, general-purpose, and supervised LLM text detector, SElected-Next-Token tRAnsformer (SENTRA). SENTRA is a Transformer-based encoder leveraging selected-next-token-probability sequences and utilizing contrastive pre-training on large amounts of unlabeled data. Our experiments on three popular public datasets across 24 domains of text demonstrate SENTRA is a general-purpose classifier that significantly outperforms popular baselines in the out-of-domain setting. |
| title | SENTRA: Selected-Next-Token Transformer for LLM Text Detection |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2509.12385 |