MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yuanfan, Zhou, Qi, Li, Chengzhengxu, Zhang, Zhaohan, Zhao, Chenxu, Ruan, Zepu, Shen, Chao, Liu, Xiaoming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
von: Li, Yuanfan, et al.
Veröffentlicht: (2025)
von: Li, Yuanfan, et al.
Veröffentlicht: (2025)
Fight Poison with Poison: Enhancing Robustness in Few-shot Machine-Generated Text Detection with Adversarial Training
von: Duan, Wenjing, et al.
Veröffentlicht: (2026)
von: Duan, Wenjing, et al.
Veröffentlicht: (2026)
GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
von: Meng, Wenlong, et al.
Veröffentlicht: (2025)
von: Meng, Wenlong, et al.
Veröffentlicht: (2025)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
von: Liu, Songyang, et al.
Veröffentlicht: (2026)
von: Liu, Songyang, et al.
Veröffentlicht: (2026)
Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID
von: Han, Xia, et al.
Veröffentlicht: (2025)
von: Han, Xia, et al.
Veröffentlicht: (2025)
Task-Agnostic Detector for Insertion-Based Backdoor Attacks
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
A Character-based Diffusion Embedding Algorithm for Enhancing the Generation Quality of Generative Linguistic Steganographic Texts
von: Chen, Yingquan, et al.
Veröffentlicht: (2025)
von: Chen, Yingquan, et al.
Veröffentlicht: (2025)
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
von: Wang, Lionel Z., et al.
Veröffentlicht: (2026)
von: Wang, Lionel Z., et al.
Veröffentlicht: (2026)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
von: Zheng, Jingyi, et al.
Veröffentlicht: (2025)
von: Zheng, Jingyi, et al.
Veröffentlicht: (2025)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
All Your Knowledge Belongs to Us: Stealing Knowledge Graphs via Reasoning APIs
von: Xi, Zhaohan
Veröffentlicht: (2025)
von: Xi, Zhaohan
Veröffentlicht: (2025)
OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models
von: Wang, Thomas, et al.
Veröffentlicht: (2025)
von: Wang, Thomas, et al.
Veröffentlicht: (2025)
StructuralSleight: Automated Jailbreak Attacks on Large Language Models Utilizing Uncommon Text-Organization Structures
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
Bileve: Securing Text Provenance in Large Language Models Against Spoofing with Bi-level Signature
von: Zhou, Tong, et al.
Veröffentlicht: (2024)
von: Zhou, Tong, et al.
Veröffentlicht: (2024)
Automated Trustworthiness Oracle Generation for Machine Learning Text Classifiers
von: Tung, Lam Nguyen, et al.
Veröffentlicht: (2024)
von: Tung, Lam Nguyen, et al.
Veröffentlicht: (2024)
HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection
von: Dai, Fangqi, et al.
Veröffentlicht: (2025)
von: Dai, Fangqi, et al.
Veröffentlicht: (2025)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge?
von: Wang, Shang, et al.
Veröffentlicht: (2024)
von: Wang, Shang, et al.
Veröffentlicht: (2024)
Adversarial Text Generation with Dynamic Contextual Perturbation
von: Waghela, Hetvi, et al.
Veröffentlicht: (2025)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2025)
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
von: Hastuti, Rochana Prih, et al.
Veröffentlicht: (2025)
von: Hastuti, Rochana Prih, et al.
Veröffentlicht: (2025)
Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models
von: Li, Haoran, et al.
Veröffentlicht: (2023)
von: Li, Haoran, et al.
Veröffentlicht: (2023)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
von: Li, Haoran, et al.
Veröffentlicht: (2024)
von: Li, Haoran, et al.
Veröffentlicht: (2024)
A General Pseudonymization Framework for Cloud-Based LLMs: Replacing Privacy Information in Controlled Text Generation
von: Hou, Shilong, et al.
Veröffentlicht: (2025)
von: Hou, Shilong, et al.
Veröffentlicht: (2025)
Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
von: Zhang, Jiankun, et al.
Veröffentlicht: (2025)
von: Zhang, Jiankun, et al.
Veröffentlicht: (2025)
TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity
von: Cao, Xi, et al.
Veröffentlicht: (2024)
von: Cao, Xi, et al.
Veröffentlicht: (2024)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
von: Pathade, Chetan
Veröffentlicht: (2025)
von: Pathade, Chetan
Veröffentlicht: (2025)
Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction
von: Zhang, Jinchuan, et al.
Veröffentlicht: (2024)
von: Zhang, Jinchuan, et al.
Veröffentlicht: (2024)
Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
von: Liu, Lijia, et al.
Veröffentlicht: (2025)
von: Liu, Lijia, et al.
Veröffentlicht: (2025)
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
von: Luo, Zhifan, et al.
Veröffentlicht: (2025)
von: Luo, Zhifan, et al.
Veröffentlicht: (2025)
Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark Remedy
von: Fu, Yu, et al.
Veröffentlicht: (2023)
von: Fu, Yu, et al.
Veröffentlicht: (2023)
Beyond Theoretical Bounds: Empirical Privacy Loss Calibration for Text Rewriting Under Local Differential Privacy
von: Li, Weijun, et al.
Veröffentlicht: (2026)
von: Li, Weijun, et al.
Veröffentlicht: (2026)
One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
von: Li, Linbao, et al.
Veröffentlicht: (2025)
von: Li, Linbao, et al.
Veröffentlicht: (2025)
LLM-based Privacy Data Augmentation Guided by Knowledge Distillation with a Distribution Tutor for Medical Text Classification
von: Song, Yiping, et al.
Veröffentlicht: (2024)
von: Song, Yiping, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
von: Li, Yuanfan, et al.
Veröffentlicht: (2025) -
Fight Poison with Poison: Enhancing Robustness in Few-shot Machine-Generated Text Detection with Adversarial Training
von: Duan, Wenjing, et al.
Veröffentlicht: (2026) -
GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
von: Meng, Wenlong, et al.
Veröffentlicht: (2025) -
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
von: Liu, Zesen, et al.
Veröffentlicht: (2024) -
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
von: Liu, Songyang, et al.
Veröffentlicht: (2026)