GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
Fuente:
arXiv
Saved in:
| Main Authors: | Meng, Wenlong, Fan, Shuguo, Wei, Chengkun, Chen, Min, Li, Yuwei, Zhang, Yuanchao, Zhang, Zhikun, Chen, Wenzhi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DC-SGD: Differentially Private SGD with Dynamic Clipping through Gradient Norm Distribution Estimation
by: Wei, Chengkun, et al.
Published: (2025)
by: Wei, Chengkun, et al.
Published: (2025)
Watermarking LLM Agent Trajectories
by: Meng, Wenlong, et al.
Published: (2026)
by: Meng, Wenlong, et al.
Published: (2026)
DWBench: Holistic Evaluation of Watermark for Dataset Copyright Auditing
by: Ren, Xiao, et al.
Published: (2026)
by: Ren, Xiao, et al.
Published: (2026)
Activation Gradient based Poisoned Sample Detection Against Backdoor Attacks
by: Yuan, Danni, et al.
Published: (2023)
by: Yuan, Danni, et al.
Published: (2023)
InferDPT: Privacy-Preserving Inference for Closed-box Large Language Model
by: Tong, Meng, et al.
Published: (2023)
by: Tong, Meng, et al.
Published: (2023)
KUDA: Knowledge Unlearning by Deviating Representation for Large Language Models
by: Fang, Ce, et al.
Published: (2026)
by: Fang, Ce, et al.
Published: (2026)
ArtistAuditor: Auditing Artist Style Pirate in Text-to-Image Generation Models
by: Du, Linkang, et al.
Published: (2025)
by: Du, Linkang, et al.
Published: (2025)
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
by: Zheng, Jingyi, et al.
Published: (2025)
by: Zheng, Jingyi, et al.
Published: (2025)
On the Vulnerability of Text Sanitization
by: Tong, Meng, et al.
Published: (2024)
by: Tong, Meng, et al.
Published: (2024)
MirrorFuzz: Leveraging LLM and Shared Bugs for Deep Learning Framework APIs Fuzzing
by: Ou, Shiwen, et al.
Published: (2025)
by: Ou, Shiwen, et al.
Published: (2025)
PrivATE: Differentially Private Average Treatment Effect Estimation for Observational Data
by: Yuan, Quan, et al.
Published: (2025)
by: Yuan, Quan, et al.
Published: (2025)
PSGraph: Differentially Private Streaming Graph Synthesis by Considering Temporal Dynamics
by: Yuan, Quan, et al.
Published: (2024)
by: Yuan, Quan, et al.
Published: (2024)
On the Credibility of Backdoor Attacks Against Object Detectors in the Physical World
by: Doan, Bao Gia, et al.
Published: (2024)
by: Doan, Bao Gia, et al.
Published: (2024)
Enhanced Privacy Leakage from Noise-Perturbed Gradients via Gradient-Guided Conditional Diffusion Models
by: Meng, Jiayang, et al.
Published: (2025)
by: Meng, Jiayang, et al.
Published: (2025)
Prompt Stealing Attacks Against Text-to-Image Generation Models
by: Shen, Xinyue, et al.
Published: (2023)
by: Shen, Xinyue, et al.
Published: (2023)
Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs
by: Xiang, Shiyu, et al.
Published: (2025)
by: Xiang, Shiyu, et al.
Published: (2025)
MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors
by: Li, Yuanfan, et al.
Published: (2026)
by: Li, Yuanfan, et al.
Published: (2026)
CENSOR: Defense Against Gradient Inversion via Orthogonal Subspace Bayesian Sampling
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
HogVul: Black-box Adversarial Code Generation Framework Against LM-based Vulnerability Detectors
by: Yang, Jingxiao, et al.
Published: (2026)
by: Yang, Jingxiao, et al.
Published: (2026)
SoK: Dataset Copyright Auditing in Machine Learning Systems
by: Du, Linkang, et al.
Published: (2024)
by: Du, Linkang, et al.
Published: (2024)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
by: Xie, Yueqi, et al.
Published: (2024)
by: Xie, Yueqi, et al.
Published: (2024)
A New Federated Learning Framework Against Gradient Inversion Attacks
by: Guo, Pengxin, et al.
Published: (2024)
by: Guo, Pengxin, et al.
Published: (2024)
Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation
by: Gao, Yilan, et al.
Published: (2026)
by: Gao, Yilan, et al.
Published: (2026)
Lifefin: Escaping Mempool Explosions in DAG-based BFT
by: Zhang, Jianting, et al.
Published: (2025)
by: Zhang, Jianting, et al.
Published: (2025)
ICSFuzz: Collision Detector Bug Discovery in Autonomous Driving Simulators
by: Fu, Weiwei, et al.
Published: (2024)
by: Fu, Weiwei, et al.
Published: (2024)
PCDM: A Diffusion-Based Data Poisoning Attack Against Federated Learning Systems
by: Sun, Wei, et al.
Published: (2026)
by: Sun, Wei, et al.
Published: (2026)
LiteUpdate: A Lightweight Framework for Updating AI-Generated Image Detectors
by: Lu, Jiajie, et al.
Published: (2025)
by: Lu, Jiajie, et al.
Published: (2025)
TextCrafter: Optimization-Calibrated Noise for Defending Against Text Embedding Inversion
by: Tang, Duoxun, et al.
Published: (2025)
by: Tang, Duoxun, et al.
Published: (2025)
Provably Robust Multi-bit Watermarking for AI-generated Text
by: Qu, Wenjie, et al.
Published: (2024)
by: Qu, Wenjie, et al.
Published: (2024)
TA3: Testing Against Adversarial Attacks on Machine Learning Models
by: Jin, Yuanzhe, et al.
Published: (2024)
by: Jin, Yuanzhe, et al.
Published: (2024)
The Gradient Puppeteer: Adversarial Domination in Gradient Leakage Attacks through Model Poisoning
by: Xiang, Kunlan, et al.
Published: (2025)
by: Xiang, Kunlan, et al.
Published: (2025)
ROAST: Risk-aware Outlier-exposure for Adversarial Selective Training of Anomaly Detectors Against Evasion Attacks
by: Elnawawy, Mohammed, et al.
Published: (2026)
by: Elnawawy, Mohammed, et al.
Published: (2026)
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
by: Li, Yuanfan, et al.
Published: (2025)
by: Li, Yuanfan, et al.
Published: (2025)
Optimal Defenses Against Gradient Reconstruction Attacks
by: Chen, Yuxiao, et al.
Published: (2024)
by: Chen, Yuxiao, et al.
Published: (2024)
Bayes-Nash Generative Privacy Against Membership Inference Attacks
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
Can't Slow me Down: Learning Robust and Hardware-Adaptive Object Detectors against Latency Attacks for Edge Devices
by: Wang, Tianyi, et al.
Published: (2024)
by: Wang, Tianyi, et al.
Published: (2024)
Automatic Jailbreaking of the Text-to-Image Generative AI Systems
by: Kim, Minseon, et al.
Published: (2024)
by: Kim, Minseon, et al.
Published: (2024)
Real-Time Trajectory Synthesis with Local Differential Privacy
by: Hu, Yujia, et al.
Published: (2024)
by: Hu, Yujia, et al.
Published: (2024)
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
by: Ranganath, Suraj, et al.
Published: (2026)
by: Ranganath, Suraj, et al.
Published: (2026)
A Model Stealing Attack Against Multi-Exit Networks
by: Pan, Li, et al.
Published: (2023)
by: Pan, Li, et al.
Published: (2023)
Similar Items
-
DC-SGD: Differentially Private SGD with Dynamic Clipping through Gradient Norm Distribution Estimation
by: Wei, Chengkun, et al.
Published: (2025) -
Watermarking LLM Agent Trajectories
by: Meng, Wenlong, et al.
Published: (2026) -
DWBench: Holistic Evaluation of Watermark for Dataset Copyright Auditing
by: Ren, Xiao, et al.
Published: (2026) -
Activation Gradient based Poisoned Sample Detection Against Backdoor Attacks
by: Yuan, Danni, et al.
Published: (2023) -
InferDPT: Privacy-Preserving Inference for Closed-box Large Language Model
by: Tong, Meng, et al.
Published: (2023)