Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Roth, Tom, Gao, Yansong, Abuadbba, Alsharif, Nepal, Surya, Liu, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DeepTaster: Adversarial Perturbation-Based Fingerprinting to Identify Proprietary Dataset Use in Deep Neural Networks
von: Park, Seonhye, et al.
Veröffentlicht: (2022)
von: Park, Seonhye, et al.
Veröffentlicht: (2022)
Adversarial Attacks Against Automated Fact-Checking: A Survey
von: Liu, Fanzhen, et al.
Veröffentlicht: (2025)
von: Liu, Fanzhen, et al.
Veröffentlicht: (2025)
Comprehensive Evaluation of Cloaking Backdoor Attacks on Object Detector in Real-World
von: Ma, Hua, et al.
Veröffentlicht: (2025)
von: Ma, Hua, et al.
Veröffentlicht: (2025)
SoK: Can Trajectory Generation Combine Privacy and Utility?
von: Buchholz, Erik, et al.
Veröffentlicht: (2024)
von: Buchholz, Erik, et al.
Veröffentlicht: (2024)
Large Language Model Adversarial Landscape Through the Lens of Attack Objectives
von: Wang, Nan, et al.
Veröffentlicht: (2025)
von: Wang, Nan, et al.
Veröffentlicht: (2025)
From Solitary Directives to Interactive Encouragement! LLM Secure Code Generation by Natural Language Prompting
von: Liu, Shigang, et al.
Veröffentlicht: (2024)
von: Liu, Shigang, et al.
Veröffentlicht: (2024)
What is the Cost of Differential Privacy for Deep Learning-Based Trajectory Generation?
von: Buchholz, Erik, et al.
Veröffentlicht: (2025)
von: Buchholz, Erik, et al.
Veröffentlicht: (2025)
Watch Out! Simple Horizontal Class Backdoor Can Trivially Evade Defense
von: Ma, Hua, et al.
Veröffentlicht: (2023)
von: Ma, Hua, et al.
Veröffentlicht: (2023)
Reversible Jump Attack to Textual Classifiers with Modification Reduction
von: Ni, Mingze, et al.
Veröffentlicht: (2024)
von: Ni, Mingze, et al.
Veröffentlicht: (2024)
Adversarially Guided Stateful Defense Against Backdoor Attacks in Federated Deep Learning
von: Ali, Hassan, et al.
Veröffentlicht: (2024)
von: Ali, Hassan, et al.
Veröffentlicht: (2024)
Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
Human Society-Inspired Approaches to Agentic AI Security: The 4C Framework
von: Abuadbba, Alsharif, et al.
Veröffentlicht: (2026)
von: Abuadbba, Alsharif, et al.
Veröffentlicht: (2026)
NADD: Amplifying Noise for Effective Diffusion-based Adversarial Purification
von: Nguyen, David D., et al.
Veröffentlicht: (2026)
von: Nguyen, David D., et al.
Veröffentlicht: (2026)
Cross-Entropy Attacks to Language Models via Rare Event Simulation
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
Can Current Detectors Catch Face-to-Voice Deepfake Attacks?
von: Nguyen, Nguyen Linh Bao, et al.
Veröffentlicht: (2025)
von: Nguyen, Nguyen Linh Bao, et al.
Veröffentlicht: (2025)
IDT: Dual-Task Adversarial Attacks for Privacy Protection
von: Faustini, Pedro, et al.
Veröffentlicht: (2024)
von: Faustini, Pedro, et al.
Veröffentlicht: (2024)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
von: More, Yash, et al.
Veröffentlicht: (2024)
von: More, Yash, et al.
Veröffentlicht: (2024)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
Contextual Chart Generation for Cyber Deception
von: Nguyen, David D., et al.
Veröffentlicht: (2024)
von: Nguyen, David D., et al.
Veröffentlicht: (2024)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models
von: Liu, Xiaoze, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2025)
Semantic Stealth: Adversarial Text Attacks on NLP Using Several Methods
von: Dey, Roopkatha, et al.
Veröffentlicht: (2024)
von: Dey, Roopkatha, et al.
Veröffentlicht: (2024)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
A Modified Word Saliency-Based Adversarial Attack on Text Classification Models
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
von: Zhang, Fangyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Fangyuan, et al.
Veröffentlicht: (2024)
An Investigation into Misuse of Java Security APIs by Large Language Models
von: Mousavi, Zahra, et al.
Veröffentlicht: (2024)
von: Mousavi, Zahra, et al.
Veröffentlicht: (2024)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Exploring Vulnerabilities and Protections in Large Language Models: A Survey
von: Liu, Frank Weizhen, et al.
Veröffentlicht: (2024)
von: Liu, Frank Weizhen, et al.
Veröffentlicht: (2024)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
Differentially Private Next-Token Prediction of Large Language Models
von: Flemings, James, et al.
Veröffentlicht: (2024)
von: Flemings, James, et al.
Veröffentlicht: (2024)
User Inference Attacks on Large Language Models
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
DeepTaster: Adversarial Perturbation-Based Fingerprinting to Identify Proprietary Dataset Use in Deep Neural Networks
von: Park, Seonhye, et al.
Veröffentlicht: (2022) -
Adversarial Attacks Against Automated Fact-Checking: A Survey
von: Liu, Fanzhen, et al.
Veröffentlicht: (2025) -
Comprehensive Evaluation of Cloaking Backdoor Attacks on Object Detector in Real-World
von: Ma, Hua, et al.
Veröffentlicht: (2025) -
SoK: Can Trajectory Generation Combine Privacy and Utility?
von: Buchholz, Erik, et al.
Veröffentlicht: (2024) -
Large Language Model Adversarial Landscape Through the Lens of Attack Objectives
von: Wang, Nan, et al.
Veröffentlicht: (2025)