Gespeichert in:
| 1. Verfasser: | Zhai, Keke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2501.00517 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recent Advances in Attack and Defense Approaches of Large Language Models
von: Cui, Jing, et al.
Veröffentlicht: (2024)
von: Cui, Jing, et al.
Veröffentlicht: (2024)
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
von: Yang, Xianglin, et al.
Veröffentlicht: (2025)
von: Yang, Xianglin, et al.
Veröffentlicht: (2025)
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
von: Tong, Haibo, et al.
Veröffentlicht: (2025)
von: Tong, Haibo, et al.
Veröffentlicht: (2025)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
Emerging Safety Attack and Defense in Federated Instruction Tuning of Large Language Models
von: Ye, Rui, et al.
Veröffentlicht: (2024)
von: Ye, Rui, et al.
Veröffentlicht: (2024)
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models
von: Chen, Yulong, et al.
Veröffentlicht: (2025)
von: Chen, Yulong, et al.
Veröffentlicht: (2025)
A Survey on Model Extraction Attacks and Defenses for Large Language Models
von: Zhao, Kaixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Kaixiang, et al.
Veröffentlicht: (2025)
Dual Defense: Enhancing Privacy and Mitigating Poisoning Attacks in Federated Learning
von: Xu, Runhua, et al.
Veröffentlicht: (2025)
von: Xu, Runhua, et al.
Veröffentlicht: (2025)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
von: Guan, Shaowei, et al.
Veröffentlicht: (2025)
von: Guan, Shaowei, et al.
Veröffentlicht: (2025)
Attacks and Defenses for Generative Diffusion Models: A Comprehensive Survey
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2024)
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2024)
Surviving the Unseen: Predictive Defense for Novel Multi-Turn Multimodal Attacks
von: You, Doohee
Veröffentlicht: (2026)
von: You, Doohee
Veröffentlicht: (2026)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
Securing Retrieval-Augmented Generation: A Taxonomy of Attacks, Defenses, and Future Directions
von: Xu, Yuming, et al.
Veröffentlicht: (2026)
von: Xu, Yuming, et al.
Veröffentlicht: (2026)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2026)
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2026)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
GUARD-SLM: Token Activation-Based Defense Against Jailbreak Attacks for Small Language Models
von: Mia, Md Jueal, et al.
Veröffentlicht: (2026)
von: Mia, Md Jueal, et al.
Veröffentlicht: (2026)
Threats, Attacks, and Defenses in Machine Unlearning: A Survey
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems
von: Balashov, Andrii, et al.
Veröffentlicht: (2025)
von: Balashov, Andrii, et al.
Veröffentlicht: (2025)
Intellectual Property in Graph-Based Machine Learning as a Service: Attacks and Defenses
von: Li, Lincan, et al.
Veröffentlicht: (2025)
von: Li, Lincan, et al.
Veröffentlicht: (2025)
Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
von: Li, Xurui, et al.
Veröffentlicht: (2025)
von: Li, Xurui, et al.
Veröffentlicht: (2025)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
von: Shang, Zhengchun, et al.
Veröffentlicht: (2025)
von: Shang, Zhengchun, et al.
Veröffentlicht: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
von: Kim, Juhee, et al.
Veröffentlicht: (2026)
von: Kim, Juhee, et al.
Veröffentlicht: (2026)
AISA: Awakening Intrinsic Safety Awareness in Large Language Models against Jailbreak Attacks
von: Song, Weiming, et al.
Veröffentlicht: (2026)
von: Song, Weiming, et al.
Veröffentlicht: (2026)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
von: Li, Xiaohu, et al.
Veröffentlicht: (2025)
von: Li, Xiaohu, et al.
Veröffentlicht: (2025)
Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models
von: Gu, Tianle, et al.
Veröffentlicht: (2024)
von: Gu, Tianle, et al.
Veröffentlicht: (2024)
Defenses Against Prompt Attacks Learn Surface Heuristics
von: Li, Shawn, et al.
Veröffentlicht: (2026)
von: Li, Shawn, et al.
Veröffentlicht: (2026)
Exploring Backdoor Attack and Defense for LLM-empowered Recommendations
von: Ning, Liangbo, et al.
Veröffentlicht: (2025)
von: Ning, Liangbo, et al.
Veröffentlicht: (2025)
Adversarial Machine Learning: Attacks, Defenses, and Open Challenges
von: Jha, Pranav K
Veröffentlicht: (2025)
von: Jha, Pranav K
Veröffentlicht: (2025)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
von: Liu, Shuyuan, et al.
Veröffentlicht: (2025)
von: Liu, Shuyuan, et al.
Veröffentlicht: (2025)
Investigating the Impact of Quantization Methods on the Safety and Reliability of Large Language Models
von: Kharinaev, Artyom, et al.
Veröffentlicht: (2025)
von: Kharinaev, Artyom, et al.
Veröffentlicht: (2025)
Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS
von: Ennaji, Sabrine, et al.
Veröffentlicht: (2025)
von: Ennaji, Sabrine, et al.
Veröffentlicht: (2025)
A General Black-box Adversarial Attack on Graph-based Fake News Detectors
von: Zhu, Peican, et al.
Veröffentlicht: (2024)
von: Zhu, Peican, et al.
Veröffentlicht: (2024)
Evaluation of Prompt Injection Defenses in Large Language Models
von: Deep, Priyal, et al.
Veröffentlicht: (2026)
von: Deep, Priyal, et al.
Veröffentlicht: (2026)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
von: Yeo, Andrew, et al.
Veröffentlicht: (2025)
von: Yeo, Andrew, et al.
Veröffentlicht: (2025)
Adversarial Attacks and Defenses on Graph-aware Large Language Models (LLMs)
von: Olatunji, Iyiola E., et al.
Veröffentlicht: (2025)
von: Olatunji, Iyiola E., et al.
Veröffentlicht: (2025)
Enhancing O-RAN Security: Evasion Attacks and Robust Defenses for Graph Reinforcement Learning-based Connection Management
von: Balakrishnan, Ravikumar, et al.
Veröffentlicht: (2024)
von: Balakrishnan, Ravikumar, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Recent Advances in Attack and Defense Approaches of Large Language Models
von: Cui, Jing, et al.
Veröffentlicht: (2024) -
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
von: Yang, Xianglin, et al.
Veröffentlicht: (2025) -
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
von: Xu, Zihao, et al.
Veröffentlicht: (2024) -
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
von: Tong, Haibo, et al.
Veröffentlicht: (2025) -
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)