Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Yize, Sadasivan, Vinu Sankar, Saberi, Mehrdad, Saha, Shoumik, Feizi, Soheil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fast Adversarial Attacks on Language Models In One GPU Minute
by: Sadasivan, Vinu Sankar, et al.
Published: (2024)
by: Sadasivan, Vinu Sankar, et al.
Published: (2024)
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing
by: Saha, Shoumik, et al.
Published: (2025)
by: Saha, Shoumik, et al.
Published: (2025)
Can AI-Generated Text be Reliably Detected?
by: Sadasivan, Vinu Sankar, et al.
Published: (2023)
by: Sadasivan, Vinu Sankar, et al.
Published: (2023)
DREW : Towards Robust Data Provenance by Leveraging Error-Controlled Watermarking
by: Saberi, Mehrdad, et al.
Published: (2024)
by: Saberi, Mehrdad, et al.
Published: (2024)
Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks
by: Saberi, Mehrdad, et al.
Published: (2023)
by: Saberi, Mehrdad, et al.
Published: (2023)
SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents
by: Saberi, Mehrdad, et al.
Published: (2026)
by: Saberi, Mehrdad, et al.
Published: (2026)
Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
by: Saha, Shoumik, et al.
Published: (2026)
by: Saha, Shoumik, et al.
Published: (2026)
DyePack: Provably Flagging Test Set Contamination in LLMs Using Backdoors
by: Cheng, Yize, et al.
Published: (2025)
by: Cheng, Yize, et al.
Published: (2025)
Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
PRIME: Prioritizing Interpretability in Failure Mode Extraction
by: Rezaei, Keivan, et al.
Published: (2023)
by: Rezaei, Keivan, et al.
Published: (2023)
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception
by: Cheng, Yize, et al.
Published: (2025)
by: Cheng, Yize, et al.
Published: (2025)
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
by: Soto, Rafael Rivera, et al.
Published: (2024)
by: Soto, Rafael Rivera, et al.
Published: (2024)
Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings
by: Zarei, Arman, et al.
Published: (2024)
by: Zarei, Arman, et al.
Published: (2024)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
by: Kaneko, Masahiro
Published: (2026)
by: Kaneko, Masahiro
Published: (2026)
PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks
by: Zha, Yiwei, et al.
Published: (2025)
by: Zha, Yiwei, et al.
Published: (2025)
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
by: Chegini, Atoosa, et al.
Published: (2026)
by: Chegini, Atoosa, et al.
Published: (2026)
Maestro: Joint Graph & Config Optimization for Reliable AI Agents
by: Wang, Wenxiao, et al.
Published: (2025)
by: Wang, Wenxiao, et al.
Published: (2025)
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
by: Zhou, Ying, et al.
Published: (2024)
by: Zhou, Ying, et al.
Published: (2024)
Understanding the Effects of Human-written Paraphrases in LLM-generated Text Detection
by: Lau, Hiu Ting, et al.
Published: (2024)
by: Lau, Hiu Ting, et al.
Published: (2024)
AdParaphrase: Paraphrase Dataset for Analyzing Linguistic Features toward Generating Attractive Ad Texts
by: Murakami, Soichiro, et al.
Published: (2025)
by: Murakami, Soichiro, et al.
Published: (2025)
A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts
by: Tripto, Nafis Irtiza, et al.
Published: (2023)
by: Tripto, Nafis Irtiza, et al.
Published: (2023)
Revisiting the Past: Data Unlearning with Model State History
by: Rezaei, Keivan, et al.
Published: (2025)
by: Rezaei, Keivan, et al.
Published: (2025)
Endor: Hardware-Friendly Sparse Format for Offloaded LLM Inference
by: Joo, Donghyeon, et al.
Published: (2024)
by: Joo, Donghyeon, et al.
Published: (2024)
AdParaphrase v2.0: Generating Attractive Ad Texts Using a Preference-Annotated Paraphrase Dataset
by: Murakami, Soichiro, et al.
Published: (2025)
by: Murakami, Soichiro, et al.
Published: (2025)
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
by: Wang, Wenxiao, et al.
Published: (2025)
by: Wang, Wenxiao, et al.
Published: (2025)
Spotting AI's Touch: Identifying LLM-Paraphrased Spans in Text
by: Li, Yafu, et al.
Published: (2024)
by: Li, Yafu, et al.
Published: (2024)
A Human Word Association based model for topic detection in social networks
by: Khadivi, Mehrdad Ranjbar, et al.
Published: (2023)
by: Khadivi, Mehrdad Ranjbar, et al.
Published: (2023)
Tool Preferences in Agentic LLMs are Unreliable
by: Faghih, Kazem, et al.
Published: (2025)
by: Faghih, Kazem, et al.
Published: (2025)
A Generative Adversarial Attack for Multilingual Text Classifiers
by: Roth, Tom, et al.
Published: (2024)
by: Roth, Tom, et al.
Published: (2024)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
by: Teja, Lekkala Sai, et al.
Published: (2025)
by: Teja, Lekkala Sai, et al.
Published: (2025)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
by: Balasubramanian, Sriram, et al.
Published: (2025)
by: Balasubramanian, Sriram, et al.
Published: (2025)
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)
by: Kumar, Aounon, et al.
Published: (2023)
Revisiting the Robustness of Watermarking to Paraphrasing Attacks
by: Rastogi, Saksham, et al.
Published: (2024)
by: Rastogi, Saksham, et al.
Published: (2024)
Paraphrase Types for Generation and Detection
by: Wahle, Jan Philip, et al.
Published: (2023)
by: Wahle, Jan Philip, et al.
Published: (2023)
Linguistically-Controlled Paraphrase Generation
by: Elgaar, Mohamed, et al.
Published: (2024)
by: Elgaar, Mohamed, et al.
Published: (2024)
VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark Models
by: Cheng, Ming, et al.
Published: (2024)
by: Cheng, Ming, et al.
Published: (2024)
LAMPAT: Low-Rank Adaption for Multilingual Paraphrasing Using Adversarial Training
by: Le, Khoi M., et al.
Published: (2024)
by: Le, Khoi M., et al.
Published: (2024)
Similar Items
-
Fast Adversarial Attacks on Language Models In One GPU Minute
by: Sadasivan, Vinu Sankar, et al.
Published: (2024) -
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
by: Sadasivan, Vinu Sankar, et al.
Published: (2025) -
Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing
by: Saha, Shoumik, et al.
Published: (2025) -
Can AI-Generated Text be Reliably Detected?
by: Sadasivan, Vinu Sankar, et al.
Published: (2023) -
DREW : Towards Robust Data Provenance by Leveraging Error-Controlled Watermarking
by: Saberi, Mehrdad, et al.
Published: (2024)