Attacks on the neural network and defense methods
Fuente:
arXiv
Saved in:
| Main Authors: | Korenev, A., Belokrylov, G., Lodonova, B., Novokhrestov, A. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ensemble of classifiers for speech evaluation
by: Belokrylov, G., et al.
Published: (2024)
by: Belokrylov, G., et al.
Published: (2024)
Watermark Overwriting Attack on StegaStamp algorithm
by: Serzhenko, I. F., et al.
Published: (2025)
by: Serzhenko, I. F., et al.
Published: (2025)
Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents
by: Shavit, Doron
Published: (2026)
by: Shavit, Doron
Published: (2026)
Per-sender neural network classifiers for email authorship validation
by: Dube, Rohit
Published: (2025)
by: Dube, Rohit
Published: (2025)
Backdoor defense, learnability and obfuscation
by: Christiano, Paul, et al.
Published: (2024)
by: Christiano, Paul, et al.
Published: (2024)
Injecting Bias into Text Classification Models using Backdoor Attacks
by: Yavuz, A. Dilara, et al.
Published: (2024)
by: Yavuz, A. Dilara, et al.
Published: (2024)
Explainable but Vulnerable: Adversarial Attacks on XAI Explanation in Cybersecurity Applications
by: Mia, Maraz, et al.
Published: (2025)
by: Mia, Maraz, et al.
Published: (2025)
FL-CLEANER: byzantine and backdoor defense by CLustering Errors of Activation maps in Non-iid fedErated leaRning
by: Ghali, Mehdi Ben, et al.
Published: (2025)
by: Ghali, Mehdi Ben, et al.
Published: (2025)
Sparsity in neural networks can improve their privacy
by: Gonon, Antoine, et al.
Published: (2023)
by: Gonon, Antoine, et al.
Published: (2023)
Attack and defense techniques in large language models: A survey and new perspectives
by: Liao, Zhiyu, et al.
Published: (2025)
by: Liao, Zhiyu, et al.
Published: (2025)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)
by: Zheng, Jingyi, et al.
Published: (2024)
Attacking Slicing Network via Side-channel Reinforcement Learning Attack
by: Shao, Wei, et al.
Published: (2024)
by: Shao, Wei, et al.
Published: (2024)
Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning
by: Song, Baogang, et al.
Published: (2025)
by: Song, Baogang, et al.
Published: (2025)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
by: Ma, Haokai, et al.
Published: (2025)
by: Ma, Haokai, et al.
Published: (2025)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
by: Kang, Mintong, et al.
Published: (2023)
by: Kang, Mintong, et al.
Published: (2023)
FIDELIS: Blockchain-Enabled Protection Against Poisoning Attacks in Federated Learning
by: Carney, Jane, et al.
Published: (2025)
by: Carney, Jane, et al.
Published: (2025)
Dynamic Target Attack
by: Xiu, Kedong, et al.
Published: (2025)
by: Xiu, Kedong, et al.
Published: (2025)
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
CAHS-Attack: CLIP-Aware Heuristic Search Attack Method for Stable Diffusion
by: Xia, Shuhan, et al.
Published: (2025)
by: Xia, Shuhan, et al.
Published: (2025)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
by: Wu, Yixin, et al.
Published: (2025)
by: Wu, Yixin, et al.
Published: (2025)
Attention Is Where You Attack
by: Srivastava, Aviral, et al.
Published: (2026)
by: Srivastava, Aviral, et al.
Published: (2026)
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
by: Cheng, Ruoxi, et al.
Published: (2024)
by: Cheng, Ruoxi, et al.
Published: (2024)
AttackER: Towards Enhancing Cyber-Attack Attribution with a Named Entity Recognition Dataset
by: Deka, Pritam, et al.
Published: (2024)
by: Deka, Pritam, et al.
Published: (2024)
PIDP-Attack: Combining Prompt Injection with Database Poisoning Attacks on Retrieval-Augmented Generation Systems
by: Wang, Haozhen, et al.
Published: (2026)
by: Wang, Haozhen, et al.
Published: (2026)
CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
by: Liu, Zesen, et al.
Published: (2025)
by: Liu, Zesen, et al.
Published: (2025)
Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success
by: Maple, Carsten, et al.
Published: (2026)
by: Maple, Carsten, et al.
Published: (2026)
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
by: Ding, Zikang, et al.
Published: (2026)
by: Ding, Zikang, et al.
Published: (2026)
Effective backdoor attack on graph neural networks in link prediction tasks
by: Dai, Jiazhu, et al.
Published: (2024)
by: Dai, Jiazhu, et al.
Published: (2024)
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025)
by: Guo, Yangyang, et al.
Published: (2025)
Security of Internet of Agents: Attacks and Countermeasures
by: Wang, Yuntao, et al.
Published: (2025)
by: Wang, Yuntao, et al.
Published: (2025)
Distillability of LLM Security Logic: Predicting Attack Success Rate of Outline Filling Attack via Ranking Regression
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
Are aligned neural networks adversarially aligned?
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization
by: Jiao, Yang, et al.
Published: (2025)
by: Jiao, Yang, et al.
Published: (2025)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
PromptLocate: Localizing Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2025)
by: Jia, Yuqi, et al.
Published: (2025)
A Survey of Attacks on Large Language Models
by: Xu, Wenrui, et al.
Published: (2025)
by: Xu, Wenrui, et al.
Published: (2025)
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
by: Zhang, Rui, et al.
Published: (2025)
by: Zhang, Rui, et al.
Published: (2025)
Model Inversion Attack against Federated Unlearning
by: Zhou, Lei, et al.
Published: (2025)
by: Zhou, Lei, et al.
Published: (2025)
Similar Items
-
Ensemble of classifiers for speech evaluation
by: Belokrylov, G., et al.
Published: (2024) -
Watermark Overwriting Attack on StegaStamp algorithm
by: Serzhenko, I. F., et al.
Published: (2025) -
Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents
by: Shavit, Doron
Published: (2026) -
Per-sender neural network classifiers for email authorship validation
by: Dube, Rohit
Published: (2025) -
Backdoor defense, learnability and obfuscation
by: Christiano, Paul, et al.
Published: (2024)