On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zesen, Cong, Tianshuo, He, Xinlei, Li, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models
von: Yi, Sibo, et al.
Veröffentlicht: (2025)
von: Yi, Sibo, et al.
Veröffentlicht: (2025)
LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models
von: Ran, Delong, et al.
Veröffentlicht: (2025)
von: Ran, Delong, et al.
Veröffentlicht: (2025)
Have You Merged My Model? On The Robustness of Large Language Model IP Protection Methods Against Model Merging
von: Cong, Tianshuo, et al.
Veröffentlicht: (2024)
von: Cong, Tianshuo, et al.
Veröffentlicht: (2024)
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation
von: Shahariar, G M, et al.
Veröffentlicht: (2024)
von: Shahariar, G M, et al.
Veröffentlicht: (2024)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
von: Zhang, Heyi, et al.
Veröffentlicht: (2025)
von: Zhang, Heyi, et al.
Veröffentlicht: (2025)
AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
von: Li, Yuexin, et al.
Veröffentlicht: (2026)
von: Li, Yuexin, et al.
Veröffentlicht: (2026)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
von: Ran, Delong, et al.
Veröffentlicht: (2024)
von: Ran, Delong, et al.
Veröffentlicht: (2024)
Downstream Trade-offs of a Family of Text Watermarks
von: Ajith, Anirudh, et al.
Veröffentlicht: (2023)
von: Ajith, Anirudh, et al.
Veröffentlicht: (2023)
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
von: Lu, Ning, et al.
Veröffentlicht: (2023)
von: Lu, Ning, et al.
Veröffentlicht: (2023)
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
Adversarial Text Purification: A Large Language Model Approach for Defense
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
Beyond Gradient and Priors in Privacy Attacks: Leveraging Pooler Layer Inputs of Language Models in Federated Learning
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
An Adversarial Perspective on Machine Unlearning for AI Safety
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
Watermarking Makes Language Models Radioactive
von: Sander, Tom, et al.
Veröffentlicht: (2024)
von: Sander, Tom, et al.
Veröffentlicht: (2024)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
von: Aldahoul, Nouar, et al.
Veröffentlicht: (2025)
von: Aldahoul, Nouar, et al.
Veröffentlicht: (2025)
Query-Based Adversarial Prompt Generation
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
Duwak: Dual Watermarks in Large Language Models
von: Zhu, Chaoyi, et al.
Veröffentlicht: (2024)
von: Zhu, Chaoyi, et al.
Veröffentlicht: (2024)
TaeBench: Improving Quality of Toxic Adversarial Examples
von: Zhu, Xuan, et al.
Veröffentlicht: (2024)
von: Zhu, Xuan, et al.
Veröffentlicht: (2024)
Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
von: Liu, Yule, et al.
Veröffentlicht: (2025)
von: Liu, Yule, et al.
Veröffentlicht: (2025)
PostMark: A Robust Blackbox Watermark for Large Language Models
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
GaussMark: A Practical Approach for Structural Watermarking of Language Models
von: Block, Adam, et al.
Veröffentlicht: (2025)
von: Block, Adam, et al.
Veröffentlicht: (2025)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
von: Teja, Lekkala Sai, et al.
Veröffentlicht: (2025)
von: Teja, Lekkala Sai, et al.
Veröffentlicht: (2025)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
Exploiting Class Probabilities for Black-box Sentence-level Attacks
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
von: Poppi, Samuele, et al.
Veröffentlicht: (2024)
von: Poppi, Samuele, et al.
Veröffentlicht: (2024)
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
von: Liu, Fuqiang, et al.
Veröffentlicht: (2024)
von: Liu, Fuqiang, et al.
Veröffentlicht: (2024)
Attack and defense techniques in large language models: A survey and new perspectives
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
von: Zheng, Jingyi, et al.
Veröffentlicht: (2025)
von: Zheng, Jingyi, et al.
Veröffentlicht: (2025)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024) -
Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models
von: Yi, Sibo, et al.
Veröffentlicht: (2025) -
LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models
von: Ran, Delong, et al.
Veröffentlicht: (2025) -
Have You Merged My Model? On The Robustness of Large Language Model IP Protection Methods Against Model Merging
von: Cong, Tianshuo, et al.
Veröffentlicht: (2024) -
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)