MGTBench: Benchmarking Machine-Generated Text Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Xinlei, Shen, Xinyue, Chen, Zeyuan, Backes, Michael, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
Prompt Stealing Attacks Against Text-to-Image Generation Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
Voice Jailbreak Attacks Against GPT-4o
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
von: Wen, Rui, et al.
Veröffentlicht: (2025)
von: Wen, Rui, et al.
Veröffentlicht: (2025)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?
von: Wen, Rui, et al.
Veröffentlicht: (2024)
von: Wen, Rui, et al.
Veröffentlicht: (2024)
Link Stealing Attacks Against Inductive Graph Neural Networks
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
von: Chen, Zeyuan, et al.
Veröffentlicht: (2026)
von: Chen, Zeyuan, et al.
Veröffentlicht: (2026)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
The Challenge of Identifying the Origin of Black-Box Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
Vera Verto: Multimodal Hijacking Attack
von: Zhang, Minxing, et al.
Veröffentlicht: (2024)
von: Zhang, Minxing, et al.
Veröffentlicht: (2024)
Provably Cost-Sensitive Adversarial Defense via Randomized Smoothing
von: Xin, Yuan, et al.
Veröffentlicht: (2023)
von: Xin, Yuan, et al.
Veröffentlicht: (2023)
Transferable Availability Poisoning Attacks
von: Liu, Yiyong, et al.
Veröffentlicht: (2023)
von: Liu, Yiyong, et al.
Veröffentlicht: (2023)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
Excessive Reasoning Attack on Reasoning LLMs
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data
von: Akkus, Atilla, et al.
Veröffentlicht: (2024)
von: Akkus, Atilla, et al.
Veröffentlicht: (2024)
Instruction Backdoor Attacks Against Customized LLMs
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability Detection
von: Risse, Niklas, et al.
Veröffentlicht: (2024)
von: Risse, Niklas, et al.
Veröffentlicht: (2024)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
von: Zhang, Heyi, et al.
Veröffentlicht: (2025)
von: Zhang, Heyi, et al.
Veröffentlicht: (2025)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
An Improved Privacy and Utility Analysis of Differentially Private SGD with Bounded Domain and Smooth Losses
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
Verification of Machine Unlearning is Fragile
von: Zhang, Binchi, et al.
Veröffentlicht: (2024)
von: Zhang, Binchi, et al.
Veröffentlicht: (2024)
Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media
von: Sun, Zhen, et al.
Veröffentlicht: (2024)
von: Sun, Zhen, et al.
Veröffentlicht: (2024)
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
von: Zhang, Yage, et al.
Veröffentlicht: (2026)
von: Zhang, Yage, et al.
Veröffentlicht: (2026)
Lossless Anti-Distillation Sampling
von: Diao, Zibo, et al.
Veröffentlicht: (2026)
von: Diao, Zibo, et al.
Veröffentlicht: (2026)
ArcGen: Generalizing Neural Backdoor Detection Across Diverse Architectures
von: Yang, Zhonghao, et al.
Veröffentlicht: (2025)
von: Yang, Zhonghao, et al.
Veröffentlicht: (2025)
k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2024)
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2024)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
Quantum Machine Learning Approaches for Coordinated Stealth Attack Detection in Distributed Generation Systems
von: Ogiesoba-Eguakun, Osasumwen Cedric, et al.
Veröffentlicht: (2025)
von: Ogiesoba-Eguakun, Osasumwen Cedric, et al.
Veröffentlicht: (2025)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
von: Wu, Yixin, et al.
Veröffentlicht: (2023)
von: Wu, Yixin, et al.
Veröffentlicht: (2023)
Unraveling the Key of Machine Learning-based Android Malware Detection
von: Liu, Jiahao, et al.
Veröffentlicht: (2024)
von: Liu, Jiahao, et al.
Veröffentlicht: (2024)
Open LLMs are Necessary for Current Private Adaptations and Outperform their Closed Alternatives
von: Hanke, Vincent, et al.
Veröffentlicht: (2024)
von: Hanke, Vincent, et al.
Veröffentlicht: (2024)
HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection
von: Sun, Danyu, et al.
Veröffentlicht: (2026)
von: Sun, Danyu, et al.
Veröffentlicht: (2026)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023) -
Prompt Stealing Attacks Against Text-to-Image Generation Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023) -
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
von: Shen, Xinyue, et al.
Veröffentlicht: (2025) -
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
von: Shen, Xinyue, et al.
Veröffentlicht: (2025) -
Voice Jailbreak Attacks Against GPT-4o
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)