Sure! Here's a short and concise title for your paper: "Contamination in Generated Text Detection Benchmarks"
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dingfelder, Philipp, Riess, Christian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
von: Schappacher-Tilp, Gudrun, et al.
Veröffentlicht: (2026)
von: Schappacher-Tilp, Gudrun, et al.
Veröffentlicht: (2026)
Estimating Exam Item Difficulty with LLMs: A Benchmark on Brazil's ENEM Corpus
von: Brant, Thiago, et al.
Veröffentlicht: (2026)
von: Brant, Thiago, et al.
Veröffentlicht: (2026)
Grandes modelos de lenguaje: de la predicción de palabras a la comprensión?
von: Gómez-Rodríguez, Carlos
Veröffentlicht: (2025)
von: Gómez-Rodríguez, Carlos
Veröffentlicht: (2025)
On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
von: Casanova, Etienne, et al.
Veröffentlicht: (2026)
von: Casanova, Etienne, et al.
Veröffentlicht: (2026)
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
von: Gautam, Sushant, et al.
Veröffentlicht: (2026)
von: Gautam, Sushant, et al.
Veröffentlicht: (2026)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code
von: Madan, Kapil
Veröffentlicht: (2025)
von: Madan, Kapil
Veröffentlicht: (2025)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free
von: Zhang, Li, et al.
Veröffentlicht: (2026)
von: Zhang, Li, et al.
Veröffentlicht: (2026)
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration
von: Guan, Xin, et al.
Veröffentlicht: (2024)
von: Guan, Xin, et al.
Veröffentlicht: (2024)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
von: Lei, Xiang, et al.
Veröffentlicht: (2025)
von: Lei, Xiang, et al.
Veröffentlicht: (2025)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
von: Zanbaghi, Shahin, et al.
Veröffentlicht: (2025)
von: Zanbaghi, Shahin, et al.
Veröffentlicht: (2025)
The Lossy Horizon: Error-Bounded Predictive Coding for Lossy Text Compression (Episode I)
von: Aghanya, Nnamdi, et al.
Veröffentlicht: (2025)
von: Aghanya, Nnamdi, et al.
Veröffentlicht: (2025)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
PairCFR: Enhancing Model Training on Paired Counterfactually Augmented Data through Contrastive Learning
von: Qiu, Xiaoqi, et al.
Veröffentlicht: (2024)
von: Qiu, Xiaoqi, et al.
Veröffentlicht: (2024)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
von: Muñoz-Ortiz, Alberto, et al.
Veröffentlicht: (2023)
von: Muñoz-Ortiz, Alberto, et al.
Veröffentlicht: (2023)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
von: Borisov, Vadim
Veröffentlicht: (2026)
von: Borisov, Vadim
Veröffentlicht: (2026)
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
von: Sela, Omer
Veröffentlicht: (2026)
von: Sela, Omer
Veröffentlicht: (2026)
Seeing Hate Differently: Hate Subspace Modeling for Culture-Aware Hate Speech Detection
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
Reviewriter: AI-Generated Instructions For Peer Review Writing
von: Su, Xiaotian, et al.
Veröffentlicht: (2025)
von: Su, Xiaotian, et al.
Veröffentlicht: (2025)
AVEC: Bootstrapping Privacy for Local LLMs
von: Gaikwad, Madhava
Veröffentlicht: (2025)
von: Gaikwad, Madhava
Veröffentlicht: (2025)
Generating Natural-Language Surgical Feedback: From Structured Representation to Domain-Grounded Evaluation
von: Nasriddinov, Firdavs, et al.
Veröffentlicht: (2025)
von: Nasriddinov, Firdavs, et al.
Veröffentlicht: (2025)
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
The Ethics Engine: A Modular Pipeline for Accessible Psychometric Assessment of Large Language Models
von: Van Clief, Jake, et al.
Veröffentlicht: (2025)
von: Van Clief, Jake, et al.
Veröffentlicht: (2025)
Learning What Matters: Probabilistic Task Selection via Mutual Information for Model Finetuning
von: Chanda, Prateek, et al.
Veröffentlicht: (2025)
von: Chanda, Prateek, et al.
Veröffentlicht: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
von: Nwokocha, Caleb Princewill
Veröffentlicht: (2022)
von: Nwokocha, Caleb Princewill
Veröffentlicht: (2022)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development
von: Szczecina, David, et al.
Veröffentlicht: (2025)
von: Szczecina, David, et al.
Veröffentlicht: (2025)
Benchmarking Deception Probes via Black-to-White Performance Boosts
von: Parrack, Avi, et al.
Veröffentlicht: (2025)
von: Parrack, Avi, et al.
Veröffentlicht: (2025)
NLP-Based Review for Toxic Comment Detection Tailored to the Chinese Cyberspace
von: Ren, Ruixing, et al.
Veröffentlicht: (2026)
von: Ren, Ruixing, et al.
Veröffentlicht: (2026)
Efficient Fine-Tuning Methods for Portuguese Question Answering: A Comparative Study of PEFT on BERTimbau and Exploratory Evaluation of Generative LLMs
von: Nina, Mariela M., et al.
Veröffentlicht: (2026)
von: Nina, Mariela M., et al.
Veröffentlicht: (2026)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
von: Wang, Yanshu, et al.
Veröffentlicht: (2026)
von: Wang, Yanshu, et al.
Veröffentlicht: (2026)
ChatGPT Based Data Augmentation for Improved Parameter-Efficient Debiasing of LLMs
von: Han, Pengrui, et al.
Veröffentlicht: (2024)
von: Han, Pengrui, et al.
Veröffentlicht: (2024)
DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
von: Zhang, Liangliang, et al.
Veröffentlicht: (2025)
von: Zhang, Liangliang, et al.
Veröffentlicht: (2025)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
Profiling German Text Simplification with Interpretable Model-Fingerprints
von: Klöser, Lars, et al.
Veröffentlicht: (2026)
von: Klöser, Lars, et al.
Veröffentlicht: (2026)
Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
von: Xu, Beining, et al.
Veröffentlicht: (2025)
von: Xu, Beining, et al.
Veröffentlicht: (2025)
A Survey of Text Watermarking in the Era of Large Language Models
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs
von: Balter, Samuel G., et al.
Veröffentlicht: (2026)
von: Balter, Samuel G., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
von: Schappacher-Tilp, Gudrun, et al.
Veröffentlicht: (2026) -
Estimating Exam Item Difficulty with LLMs: A Benchmark on Brazil's ENEM Corpus
von: Brant, Thiago, et al.
Veröffentlicht: (2026) -
Grandes modelos de lenguaje: de la predicción de palabras a la comprensión?
von: Gómez-Rodríguez, Carlos
Veröffentlicht: (2025) -
On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
von: Casanova, Etienne, et al.
Veröffentlicht: (2026) -
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
von: Gautam, Sushant, et al.
Veröffentlicht: (2026)