Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Hao, Kong, Jiawei, Zhuang, Tianqu, Qiu, Yixiang, Gao, Kuofeng, Chen, Bin, Xia, Shu-Tao, Wang, Yaowei, Zhang, Min |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization
by: Fang, Hao, et al.
Published: (2026)
by: Fang, Hao, et al.
Published: (2026)
Towards Distillation-Resistant Large Language Models: An Information-Theoretic Perspective
by: Fang, Hao, et al.
Published: (2026)
by: Fang, Hao, et al.
Published: (2026)
Rank Matters: Understanding and Defending Model Inversion Attacks via Low-Rank Feature Filtering
by: Yu, Hongyao, et al.
Published: (2024)
by: Yu, Hongyao, et al.
Published: (2024)
MIBench: A Comprehensive Framework for Benchmarking Model Inversion Attack and Defense
by: Qiu, Yixiang, et al.
Published: (2024)
by: Qiu, Yixiang, et al.
Published: (2024)
Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
by: Kong, Jiawei, et al.
Published: (2025)
by: Kong, Jiawei, et al.
Published: (2025)
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
by: Soto, Rafael Rivera, et al.
Published: (2024)
by: Soto, Rafael Rivera, et al.
Published: (2024)
Retrievals Can Be Detrimental: Unveiling the Backdoor Vulnerability of Retrieval-Augmented Diffusion Models
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks
by: Zha, Yiwei, et al.
Published: (2025)
by: Zha, Yiwei, et al.
Published: (2025)
ICAS: Detecting Training Data from Autoregressive Image Generative Models
by: Yu, Hongyao, et al.
Published: (2025)
by: Yu, Hongyao, et al.
Published: (2025)
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
by: Ranganath, Suraj, et al.
Published: (2026)
by: Ranganath, Suraj, et al.
Published: (2026)
Privacy Leakage on DNNs: A Survey of Model Inversion Attacks and Defenses
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
A Closer Look at GAN Priors: Exploiting Intermediate Features for Enhanced Model Inversion Attacks
by: Qiu, Yixiang, et al.
Published: (2024)
by: Qiu, Yixiang, et al.
Published: (2024)
Your Classifier Can Be Secretly a Likelihood-Based OOD Detector
by: Burapacheep, Jirayu, et al.
Published: (2024)
by: Burapacheep, Jirayu, et al.
Published: (2024)
Your AI-Generated Image Detector Can Secretly Achieve SOTA Accuracy, If Calibrated
by: Yang, Muli, et al.
Published: (2026)
by: Yang, Muli, et al.
Published: (2026)
Neural Antidote: Class-Wise Prompt Tuning for Purifying Backdoors in CLIP
by: Kong, Jiawei, et al.
Published: (2025)
by: Kong, Jiawei, et al.
Published: (2025)
Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text
by: Cheng, Yize, et al.
Published: (2025)
by: Cheng, Yize, et al.
Published: (2025)
Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations
by: Liu, Haitong, et al.
Published: (2025)
by: Liu, Haitong, et al.
Published: (2025)
Denial-of-Service Poisoning Attacks against Large Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP
by: Bai, Jiawang, et al.
Published: (2023)
by: Bai, Jiawang, et al.
Published: (2023)
Towards Dataset Copyright Evasion Attack against Personalized Text-to-Image Diffusion Models
by: Gao, Kuofeng, et al.
Published: (2025)
by: Gao, Kuofeng, et al.
Published: (2025)
EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts
by: Han, Yucheng, et al.
Published: (2024)
by: Han, Yucheng, et al.
Published: (2024)
Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
Revisiting the Privacy Risks of Split Inference: A GAN-Based Data Reconstruction Attack via Progressive Feature Optimization
by: Qiu, Yixiang, et al.
Published: (2025)
by: Qiu, Yixiang, et al.
Published: (2025)
Understanding the Effects of Human-written Paraphrases in LLM-generated Text Detection
by: Lau, Hiu Ting, et al.
Published: (2024)
by: Lau, Hiu Ting, et al.
Published: (2024)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
by: Kaneko, Masahiro
Published: (2026)
by: Kaneko, Masahiro
Published: (2026)
Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
by: Wang, Mingyi, et al.
Published: (2026)
by: Wang, Mingyi, et al.
Published: (2026)
Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
by: Sun, Shuoyang, et al.
Published: (2026)
by: Sun, Shuoyang, et al.
Published: (2026)
FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
Editable-DeepSC: Reliable Cross-Modal Semantic Communications for Facial Editing
by: Chen, Bin, et al.
Published: (2024)
by: Chen, Bin, et al.
Published: (2024)
TRACE: Your Diffusion Model is Secretly an Instance Edge Detector
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive Training
by: Wu, Yunshu, et al.
Published: (2024)
by: Wu, Yunshu, et al.
Published: (2024)
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator
by: Chen, Zhuotong, et al.
Published: (2025)
by: Chen, Zhuotong, et al.
Published: (2025)
Paraphrasing Attack Resilience of Various AI-Generated Text Detection Methods
by: Shportko, Andrii, et al.
Published: (2026)
by: Shportko, Andrii, et al.
Published: (2026)
Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transformers
by: Yang, Sheng, et al.
Published: (2024)
by: Yang, Sheng, et al.
Published: (2024)
RedacBench: Can AI Erase Your Secrets?
by: Jeon, Hyunjun, et al.
Published: (2026)
by: Jeon, Hyunjun, et al.
Published: (2026)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
by: Wang, Tianchun, et al.
Published: (2024)
by: Wang, Tianchun, et al.
Published: (2024)
The Erosion of LLM Signatures: Can We Still Distinguish Human and LLM-Generated Scientific Ideas After Iterative Paraphrasing?
by: Shahriar, Sadat, et al.
Published: (2025)
by: Shahriar, Sadat, et al.
Published: (2025)
Robust Multi-bit Text Watermark with LLM-based Paraphrasers
by: Xu, Xiaojun, et al.
Published: (2024)
by: Xu, Xiaojun, et al.
Published: (2024)
Similar Items
-
Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization
by: Fang, Hao, et al.
Published: (2026) -
Towards Distillation-Resistant Large Language Models: An Information-Theoretic Perspective
by: Fang, Hao, et al.
Published: (2026) -
Rank Matters: Understanding and Defending Model Inversion Attacks via Low-Rank Feature Filtering
by: Yu, Hongyao, et al.
Published: (2024) -
MIBench: A Comprehensive Framework for Benchmarking Model Inversion Attack and Defense
by: Qiu, Yixiang, et al.
Published: (2024) -
Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
by: Kong, Jiawei, et al.
Published: (2025)