TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
Fuente:
arXiv
Guardado en:
| Autores principales: | Zheng, Jingyi, Wang, Junfeng, Sun, Zhen, Dong, Wenhan, Liu, Yule, He, Xinlei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
por: Zhou, Ying, et al.
Publicado: (2024)
por: Zhou, Ying, et al.
Publicado: (2024)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
por: Gu, Yongtong, et al.
Publicado: (2026)
por: Gu, Yongtong, et al.
Publicado: (2026)
AuthorMist: Evading AI Text Detectors with Reinforcement Learning
por: David, Isaac, et al.
Publicado: (2025)
por: David, Isaac, et al.
Publicado: (2025)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
por: Liu, Zesen, et al.
Publicado: (2024)
por: Liu, Zesen, et al.
Publicado: (2024)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
por: Luo, Zeren, et al.
Publicado: (2025)
por: Luo, Zeren, et al.
Publicado: (2025)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
por: Zheng, Jingyi, et al.
Publicado: (2024)
por: Zheng, Jingyi, et al.
Publicado: (2024)
Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media
por: Sun, Zhen, et al.
Publicado: (2024)
por: Sun, Zhen, et al.
Publicado: (2024)
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
por: Sun, Zhen, et al.
Publicado: (2024)
por: Sun, Zhen, et al.
Publicado: (2024)
MGTBench: Benchmarking Machine-Generated Text Detection
por: He, Xinlei, et al.
Publicado: (2023)
por: He, Xinlei, et al.
Publicado: (2023)
Quantized Delta Weight Is Safety Keeper
por: Liu, Yule, et al.
Publicado: (2024)
por: Liu, Yule, et al.
Publicado: (2024)
MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors
por: Li, Yuanfan, et al.
Publicado: (2026)
por: Li, Yuanfan, et al.
Publicado: (2026)
Stego Battlefield: Evaluating Image Steganography Attacks and Steganalysis Defenses
por: Sun, Zhen, et al.
Publicado: (2026)
por: Sun, Zhen, et al.
Publicado: (2026)
Tarallo: Evading Behavioral Malware Detectors in the Problem Space
por: Digregorio, Gabriele, et al.
Publicado: (2025)
por: Digregorio, Gabriele, et al.
Publicado: (2025)
Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
por: Zhang, Zongmin, et al.
Publicado: (2025)
por: Zhang, Zongmin, et al.
Publicado: (2025)
Attacks on Approximate Caches in Text-to-Image Diffusion Models
por: Sun, Desen, et al.
Publicado: (2025)
por: Sun, Desen, et al.
Publicado: (2025)
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
por: Peng, Zifan, et al.
Publicado: (2025)
por: Peng, Zifan, et al.
Publicado: (2025)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
por: Zhang, Heyi, et al.
Publicado: (2025)
por: Zhang, Heyi, et al.
Publicado: (2025)
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
por: Ranganath, Suraj, et al.
Publicado: (2026)
por: Ranganath, Suraj, et al.
Publicado: (2026)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
por: Yi, Sibo, et al.
Publicado: (2024)
por: Yi, Sibo, et al.
Publicado: (2024)
Asymmetric Bias in Text-to-Image Generation with Adversarial Attacks
por: Shahgir, Haz Sameen, et al.
Publicado: (2023)
por: Shahgir, Haz Sameen, et al.
Publicado: (2023)
"To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
por: Sun, Zhen, et al.
Publicado: (2025)
por: Sun, Zhen, et al.
Publicado: (2025)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
por: Wang, Tianchun, et al.
Publicado: (2024)
por: Wang, Tianchun, et al.
Publicado: (2024)
Privacy-Preserving Federated Learning via Homomorphic Adversarial Networks
por: Dong, Wenhan, et al.
Publicado: (2024)
por: Dong, Wenhan, et al.
Publicado: (2024)
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
por: Zhong, Zhiyuan, et al.
Publicado: (2025)
por: Zhong, Zhiyuan, et al.
Publicado: (2025)
Global BGP Attacks that Evade Route Monitoring
por: Birge-Lee, Henry, et al.
Publicado: (2024)
por: Birge-Lee, Henry, et al.
Publicado: (2024)
Explainable and Transferable Adversarial Attack for ML-Based Network Intrusion Detectors
por: Zhang, Hangsheng, et al.
Publicado: (2024)
por: Zhang, Hangsheng, et al.
Publicado: (2024)
GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
por: Meng, Wenlong, et al.
Publicado: (2025)
por: Meng, Wenlong, et al.
Publicado: (2025)
EvadeDroid: A Practical Evasion Attack on Machine Learning for Black-box Android Malware Detection
por: Bostani, Hamid, et al.
Publicado: (2021)
por: Bostani, Hamid, et al.
Publicado: (2021)
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
por: Li, Yuanfan, et al.
Publicado: (2025)
por: Li, Yuanfan, et al.
Publicado: (2025)
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models
por: Liu, Yule, et al.
Publicado: (2026)
por: Liu, Yule, et al.
Publicado: (2026)
Combinational Backdoor Attack against Customized Text-to-Image Models
por: Jiang, Wenbo, et al.
Publicado: (2024)
por: Jiang, Wenbo, et al.
Publicado: (2024)
Provably Robust Multi-bit Watermarking for AI-generated Text
por: Qu, Wenjie, et al.
Publicado: (2024)
por: Qu, Wenjie, et al.
Publicado: (2024)
FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts
por: Zhang, Ziyi, et al.
Publicado: (2025)
por: Zhang, Ziyi, et al.
Publicado: (2025)
Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models
por: Shan, Shawn, et al.
Publicado: (2023)
por: Shan, Shawn, et al.
Publicado: (2023)
$PC^2$: Politically Controversial Content Generation via Jailbreaking Attacks on GPT-based Text-to-Image Models
por: Choi, Wonwoo, et al.
Publicado: (2026)
por: Choi, Wonwoo, et al.
Publicado: (2026)
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
por: Liu, Yule, et al.
Publicado: (2025)
por: Liu, Yule, et al.
Publicado: (2025)
BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting
por: Qiu, Huming, et al.
Publicado: (2023)
por: Qiu, Huming, et al.
Publicado: (2023)
DMTG: A Human-Like Mouse Trajectory Generation Bot Based on Entropy-Controlled Diffusion Networks
por: Liu, Jiahua, et al.
Publicado: (2024)
por: Liu, Jiahua, et al.
Publicado: (2024)
Comprehensive Evaluation of Cloaking Backdoor Attacks on Object Detector in Real-World
por: Ma, Hua, et al.
Publicado: (2025)
por: Ma, Hua, et al.
Publicado: (2025)
RefineRAG: Word-Level Poisoning Attacks via Retriever-Guided Text Refinement
por: Wang, Ziye, et al.
Publicado: (2026)
por: Wang, Ziye, et al.
Publicado: (2026)
Ejemplares similares
-
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
por: Zhou, Ying, et al.
Publicado: (2024) -
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
por: Gu, Yongtong, et al.
Publicado: (2026) -
AuthorMist: Evading AI Text Detectors with Reinforcement Learning
por: David, Isaac, et al.
Publicado: (2025) -
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
por: Liu, Zesen, et al.
Publicado: (2024) -
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
por: Luo, Zeren, et al.
Publicado: (2025)