How Sampling Affects the Detectability of Machine-written texts: A Comprehensive Study
Fuente:
arXiv
Saved in:
| Main Authors: | Dubois, Matthieu, Yvon, François, Piantanida, Pablo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOSAIC: Multiple Observers Spotting AI Content
by: Dubois, Matthieu, et al.
Published: (2024)
by: Dubois, Matthieu, et al.
Published: (2024)
Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset
by: Lerner, Paul, et al.
Published: (2025)
by: Lerner, Paul, et al.
Published: (2025)
Prompting LLMs: Length Control for Isometric Machine Translation
by: Javorský, Dávid, et al.
Published: (2025)
by: Javorský, Dávid, et al.
Published: (2025)
Investigating Length Issues in Document-level Machine Translation
by: Peng, Ziqian, et al.
Published: (2024)
by: Peng, Ziqian, et al.
Published: (2024)
Retrieving Examples from Memory for Retrieval Augmented Neural Machine Translation: A Systematic Comparison
by: Bouthors, Maxime, et al.
Published: (2024)
by: Bouthors, Maxime, et al.
Published: (2024)
Understanding In-Context Machine Translation for Low-Resource Languages: A Case Study on Manchu
by: Pei, Renhao, et al.
Published: (2025)
by: Pei, Renhao, et al.
Published: (2025)
Rainproof: An Umbrella To Shield Text Generators From Out-Of-Distribution Data
by: Darrin, Maxime, et al.
Published: (2022)
by: Darrin, Maxime, et al.
Published: (2022)
Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data
by: Bouthors, Maxime, et al.
Published: (2025)
by: Bouthors, Maxime, et al.
Published: (2025)
AdaptBPE: From General Purpose to Specialized Tokenizers
by: Liyanage, Vijini, et al.
Published: (2026)
by: Liyanage, Vijini, et al.
Published: (2026)
How Programming Concepts and Neurons Are Shared in Code Language Models
by: Kargaran, Amir Hossein, et al.
Published: (2025)
by: Kargaran, Amir Hossein, et al.
Published: (2025)
MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines
by: Javorský, Dávid, et al.
Published: (2025)
by: Javorský, Dávid, et al.
Published: (2025)
Unmasking the Imposters: How Censorship and Domain Adaptation Affect the Detection of Machine-Generated Tweets
by: Tuck, Bryan E., et al.
Published: (2024)
by: Tuck, Bryan E., et al.
Published: (2024)
Optimizing example selection for retrieval-augmented machine translation with translation memories
by: Bouthors, Maxime, et al.
Published: (2024)
by: Bouthors, Maxime, et al.
Published: (2024)
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
by: Labat, Léo, et al.
Published: (2026)
by: Labat, Léo, et al.
Published: (2026)
GlotScript: A Resource and Tool for Low Resource Writing System Identification
by: Kargaran, Amir Hossein, et al.
Published: (2023)
by: Kargaran, Amir Hossein, et al.
Published: (2023)
GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews
by: Darrin, Maxime, et al.
Published: (2024)
by: Darrin, Maxime, et al.
Published: (2024)
MaskLID: Code-Switching Language Identification through Iterative Masking
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
On the Entity-Level Alignment in Crosslingual Consistency
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
Collaborative Rational Speech Act: Pragmatic Reasoning for Multi-Turn Dialog
by: Estienne, Lautaro, et al.
Published: (2025)
by: Estienne, Lautaro, et al.
Published: (2025)
Understanding the Effects of Human-written Paraphrases in LLM-generated Text Detection
by: Lau, Hiu Ting, et al.
Published: (2024)
by: Lau, Hiu Ting, et al.
Published: (2024)
Using Machine Learning to Distinguish Human-written from Machine-generated Creative Fiction
by: McGlinchey, Andrea Cristina, et al.
Published: (2024)
by: McGlinchey, Andrea Cristina, et al.
Published: (2024)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
Unsupervised Layer-wise Score Aggregation for Textual OOD Detection
by: Darrin, Maxime, et al.
Published: (2023)
by: Darrin, Maxime, et al.
Published: (2023)
How do we measure privacy in text? A survey of text anonymization metrics
by: Ren, Yaxuan, et al.
Published: (2025)
by: Ren, Yaxuan, et al.
Published: (2025)
How Transliterations Improve Crosslingual Alignment
by: Liu, Yihong, et al.
Published: (2024)
by: Liu, Yihong, et al.
Published: (2024)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
by: Darrin, Maxime, et al.
Published: (2024)
by: Darrin, Maxime, et al.
Published: (2024)
GlotLID: Language Identification for Low-Resource Languages
by: Kargaran, Amir Hossein, et al.
Published: (2023)
by: Kargaran, Amir Hossein, et al.
Published: (2023)
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
by: Weck, Benno, et al.
Published: (2026)
by: Weck, Benno, et al.
Published: (2026)
$(RSA)^2$: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
by: Piano, Cesare Spinoso-Di, et al.
Published: (2025)
by: Piano, Cesare Spinoso-Di, et al.
Published: (2025)
How to predict creativity ratings from written narratives: A comparison of co-occurrence and textual forma mentis networks
by: Passaro, Roberto, et al.
Published: (2026)
by: Passaro, Roberto, et al.
Published: (2026)
Detection of Somali-written Fake News and Toxic Messages on the Social Media Using Transformer-based Language Models
by: Mohamed, Muhidin A., et al.
Published: (2025)
by: Mohamed, Muhidin A., et al.
Published: (2025)
LLMs' Reading Comprehension Is Affected by Parametric Knowledge and Struggles with Hypothetical Statements
by: Basmov, Victoria, et al.
Published: (2024)
by: Basmov, Victoria, et al.
Published: (2024)
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
How does Burrows' Delta work on medieval Chinese poetic texts?
by: Orekhov, Boris
Published: (2024)
by: Orekhov, Boris
Published: (2024)
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
by: Kargaran, Amir Hossein, et al.
Published: (2026)
by: Kargaran, Amir Hossein, et al.
Published: (2026)
How Well Do Large Reasoning Models Translate? A Comprehensive Evaluation for Multi-Domain Machine Translation
by: Ye, Yongshi, et al.
Published: (2025)
by: Ye, Yongshi, et al.
Published: (2025)
How Does the Disclosure of AI Assistance Affect the Perceptions of Writing?
by: Li, Zhuoyan, et al.
Published: (2024)
by: Li, Zhuoyan, et al.
Published: (2024)
How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations
by: Takenami, Yoshiki, et al.
Published: (2025)
by: Takenami, Yoshiki, et al.
Published: (2025)
How Does Quantization Affect Multilingual LLMs?
by: Marchisio, Kelly, et al.
Published: (2024)
by: Marchisio, Kelly, et al.
Published: (2024)
Detecting text level intellectual influence with knowledge graph embeddings
by: Li, Lucian, et al.
Published: (2024)
by: Li, Lucian, et al.
Published: (2024)
Similar Items
-
MOSAIC: Multiple Observers Spotting AI Content
by: Dubois, Matthieu, et al.
Published: (2024) -
Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset
by: Lerner, Paul, et al.
Published: (2025) -
Prompting LLMs: Length Control for Isometric Machine Translation
by: Javorský, Dávid, et al.
Published: (2025) -
Investigating Length Issues in Document-level Machine Translation
by: Peng, Ziqian, et al.
Published: (2024) -
Retrieving Examples from Memory for Retrieval Augmented Neural Machine Translation: A Systematic Comparison
by: Bouthors, Maxime, et al.
Published: (2024)