AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
Fuente:
arXiv
Saved in:
| Main Authors: | Milička, Jiří, Marklová, Anna, Cvrček, Václav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmark of stylistic variation in LLM-generated texts
by: Milička, Jiří, et al.
Published: (2025)
by: Milička, Jiří, et al.
Published: (2025)
Humans can learn to detect AI-generated texts, or at least learn when they can't
by: Milička, Jiří, et al.
Published: (2025)
by: Milička, Jiří, et al.
Published: (2025)
The author is dead, but what if they never lived? A reception experiment on Czech AI- and human-authored poetry
by: Marklová, Anna, et al.
Published: (2025)
by: Marklová, Anna, et al.
Published: (2025)
Iconicity in Large Language Models
by: Marklová, Anna, et al.
Published: (2025)
by: Marklová, Anna, et al.
Published: (2025)
Sydney Telling Fables on AI and Humans: A Corpus Tracing Memetic Transfer of Persona between LLMs
by: Milička, Jiří, et al.
Published: (2026)
by: Milička, Jiří, et al.
Published: (2026)
Hope Speech Detection in Social Media English Corpora: Performance of Traditional and Transformer Models
by: Ramos, Luis, et al.
Published: (2025)
by: Ramos, Luis, et al.
Published: (2025)
Multilingual and Explainable Text Detoxification with Parallel Corpora
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
Attributing Culture-Conditioned Generations to Pretraining Corpora
by: Li, Huihan, et al.
Published: (2024)
by: Li, Huihan, et al.
Published: (2024)
Discovering Multi-Scale Semantic Structure in Text Corpora Using Density-Based Trees and LLM Embeddings
by: Haschka, Thomas, et al.
Published: (2025)
by: Haschka, Thomas, et al.
Published: (2025)
Guylingo: The Republic of Guyana Creole Corpora
by: Clarke, Christopher, et al.
Published: (2024)
by: Clarke, Christopher, et al.
Published: (2024)
Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement
by: Kersting, Nicholas S., et al.
Published: (2026)
by: Kersting, Nicholas S., et al.
Published: (2026)
The Cross-Lingual Cost: Retrieval Biases in RAG over Arabic-English Corpora
by: Amiraz, Chen, et al.
Published: (2025)
by: Amiraz, Chen, et al.
Published: (2025)
EmbGen: Teaching with Reassembled Corpora
by: Lenin, Arun K, et al.
Published: (2026)
by: Lenin, Arun K, et al.
Published: (2026)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
Theoretical and Methodological Framework for Studying Texts Produced by Large Language Models
by: Milička, Jiří
Published: (2024)
by: Milička, Jiří
Published: (2024)
Enhancing Document-Level Machine Translation via Filtered Synthetic Corpora and Two-Stage LLM Adaptation
by: Kim, Ireh, et al.
Published: (2026)
by: Kim, Ireh, et al.
Published: (2026)
Rethinking KenLM: Good and Bad Model Ensembles for Efficient Text Quality Filtering in Large Web Corpora
by: Kim, Yungi, et al.
Published: (2024)
by: Kim, Yungi, et al.
Published: (2024)
AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese Corpora
by: Zhou, Zhihan, et al.
Published: (2025)
by: Zhou, Zhihan, et al.
Published: (2025)
Grounding Synthetic Data Evaluations of Language Models in Unsupervised Document Corpora
by: Majurski, Michael, et al.
Published: (2025)
by: Majurski, Michael, et al.
Published: (2025)
A First Context-Free Grammar Applied to Nawatl Corpora Augmentation
by: Guzmán-Landa, Juan-José, et al.
Published: (2025)
by: Guzmán-Landa, Juan-José, et al.
Published: (2025)
Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora
by: Hennara, Khalil, et al.
Published: (2025)
by: Hennara, Khalil, et al.
Published: (2025)
Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
by: Artemova, Ekaterina, et al.
Published: (2025)
by: Artemova, Ekaterina, et al.
Published: (2025)
Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora
by: Abbas, Chaymaa, et al.
Published: (2026)
by: Abbas, Chaymaa, et al.
Published: (2026)
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
by: Basu, Kinjal, et al.
Published: (2024)
by: Basu, Kinjal, et al.
Published: (2024)
A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
by: Xu, Yuemei, et al.
Published: (2024)
by: Xu, Yuemei, et al.
Published: (2024)
MegaMath: Pushing the Limits of Open Math Corpora
by: Zhou, Fan, et al.
Published: (2025)
by: Zhou, Fan, et al.
Published: (2025)
From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
by: Shen, Yingli, et al.
Published: (2025)
by: Shen, Yingli, et al.
Published: (2025)
Mitigating Stylistic Biases of Machine Translation Systems via Monolingual Corpora Only
by: Gao, Xuanqi, et al.
Published: (2025)
by: Gao, Xuanqi, et al.
Published: (2025)
SaudiBERT: A Large Language Model Pretrained on Saudi Dialect Corpora
by: Qarah, Faisal
Published: (2024)
by: Qarah, Faisal
Published: (2024)
Bottom-Up and Top-Down Analysis of Values, Agendas, and Observations in Corpora and LLMs
by: Friedman, Scott E., et al.
Published: (2024)
by: Friedman, Scott E., et al.
Published: (2024)
GhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian Languages
by: Gyamfi, Lawrence Adu, et al.
Published: (2026)
by: Gyamfi, Lawrence Adu, et al.
Published: (2026)
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
by: Bhattacharjee, Soham, et al.
Published: (2025)
by: Bhattacharjee, Soham, et al.
Published: (2025)
Classification of Human- and AI-Generated Texts for English, French, German, and Spanish
by: Schaaff, Kristina, et al.
Published: (2023)
by: Schaaff, Kristina, et al.
Published: (2023)
Simple stochastic processes behind Menzerath's Law
by: Milička, Jiří
Published: (2024)
by: Milička, Jiří
Published: (2024)
Preference Consistency Matters: Enhancing Preference Learning in Language Models with Automated Self-Curation of Training Corpora
by: Lee, JoonHo, et al.
Published: (2024)
by: Lee, JoonHo, et al.
Published: (2024)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
by: Hilasaca, Kenji, et al.
Published: (2026)
by: Hilasaca, Kenji, et al.
Published: (2026)
Semi-automated Fact-checking in Portuguese: Corpora Enrichment using Retrieval with Claim extraction
by: Gomes, Juliana Resplande Sant'anna, et al.
Published: (2025)
by: Gomes, Juliana Resplande Sant'anna, et al.
Published: (2025)
OrgForge: A Multi-Agent Simulation Framework for Verifiable Synthetic Corporate Corpora
by: Flynt, Jeffrey
Published: (2026)
by: Flynt, Jeffrey
Published: (2026)
On the Effectiveness of LLM-Specific Fine-Tuning for Detecting AI-Generated Text
by: Gromadzki, Michał, et al.
Published: (2026)
by: Gromadzki, Michał, et al.
Published: (2026)
AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora
by: Bai, Jiaxin, et al.
Published: (2025)
by: Bai, Jiaxin, et al.
Published: (2025)
Similar Items
-
Benchmark of stylistic variation in LLM-generated texts
by: Milička, Jiří, et al.
Published: (2025) -
Humans can learn to detect AI-generated texts, or at least learn when they can't
by: Milička, Jiří, et al.
Published: (2025) -
The author is dead, but what if they never lived? A reception experiment on Czech AI- and human-authored poetry
by: Marklová, Anna, et al.
Published: (2025) -
Iconicity in Large Language Models
by: Marklová, Anna, et al.
Published: (2025) -
Sydney Telling Fables on AI and Humans: A Corpus Tracing Memetic Transfer of Persona between LLMs
by: Milička, Jiří, et al.
Published: (2026)