Gespeichert in:
| Hauptverfasser: | Su, Guangzhi, Huang, Shuchang, Ke, Yutong, Liu, Zhuohang, Qian, Long, Huang, Kaizhu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2510.26830 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Defending against Jailbreak through Early Exit Generation of Large Language Models
von: Zhao, Chongwen, et al.
Veröffentlicht: (2024)
von: Zhao, Chongwen, et al.
Veröffentlicht: (2024)
KinGuard: Hierarchical Kinship-Aware Fingerprinting to Defend Against Large Language Model Stealing
von: Xu, Zhenhua, et al.
Veröffentlicht: (2026)
von: Xu, Zhenhua, et al.
Veröffentlicht: (2026)
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
Recent Advances in Attack and Defense Approaches of Large Language Models
von: Cui, Jing, et al.
Veröffentlicht: (2024)
von: Cui, Jing, et al.
Veröffentlicht: (2024)
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
von: Zhang, Shenyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shenyi, et al.
Veröffentlicht: (2025)
PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification
von: Gong, Guangyu, et al.
Veröffentlicht: (2026)
von: Gong, Guangyu, et al.
Veröffentlicht: (2026)
A Game Between the Defender and the Attacker for Trigger-based Black-box Model Watermarking
von: Huang, Chaoyue, et al.
Veröffentlicht: (2025)
von: Huang, Chaoyue, et al.
Veröffentlicht: (2025)
Enhanced Privacy Leakage from Noise-Perturbed Gradients via Gradient-Guided Conditional Diffusion Models
von: Meng, Jiayang, et al.
Veröffentlicht: (2025)
von: Meng, Jiayang, et al.
Veröffentlicht: (2025)
AGNNCert: Defending Graph Neural Networks against Arbitrary Perturbations with Deterministic Certification
von: Li, Jiate, et al.
Veröffentlicht: (2025)
von: Li, Jiate, et al.
Veröffentlicht: (2025)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Why Does Differential Privacy with Large Epsilon Defend Against Practical Membership Inference Attacks?
von: Lowy, Andrew, et al.
Veröffentlicht: (2024)
von: Lowy, Andrew, et al.
Veröffentlicht: (2024)
Invariant Aggregator for Defending against Federated Backdoor Attacks
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2022)
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2022)
TraceGuard: Process-Guided Firewall against Reasoning Backdoors in Large Language Models
von: Guo, Zhen, et al.
Veröffentlicht: (2026)
von: Guo, Zhen, et al.
Veröffentlicht: (2026)
TextCrafter: Optimization-Calibrated Noise for Defending Against Text Embedding Inversion
von: Tang, Duoxun, et al.
Veröffentlicht: (2025)
von: Tang, Duoxun, et al.
Veröffentlicht: (2025)
SmartGuard: Leveraging Large Language Models for Network Attack Detection through Audit Log Analysis and Summarization
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
DeepGuard: Secure Code Generation via Multi-Layer Semantic Aggregation
von: Huang, Li, et al.
Veröffentlicht: (2026)
von: Huang, Li, et al.
Veröffentlicht: (2026)
Large Language Models are Autonomous Cyber Defenders
von: Castro, Sebastián R., et al.
Veröffentlicht: (2025)
von: Castro, Sebastián R., et al.
Veröffentlicht: (2025)
Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
von: Lai, Songning, et al.
Veröffentlicht: (2024)
von: Lai, Songning, et al.
Veröffentlicht: (2024)
Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
von: Wong, Ryan, et al.
Veröffentlicht: (2025)
von: Wong, Ryan, et al.
Veröffentlicht: (2025)
Defending Large Language Models Against Attacks With Residual Stream Activation Analysis
von: Kawasaki, Amelia, et al.
Veröffentlicht: (2024)
von: Kawasaki, Amelia, et al.
Veröffentlicht: (2024)
Fight Perturbations with Perturbations: Defending Adversarial Attacks via Neuron Influence
von: Chen, Ruoxi, et al.
Veröffentlicht: (2021)
von: Chen, Ruoxi, et al.
Veröffentlicht: (2021)
Smooth Sensitivity for Geo-Privacy
von: Liang, Yuting, et al.
Veröffentlicht: (2024)
von: Liang, Yuting, et al.
Veröffentlicht: (2024)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
BitAbuse: A Dataset of Visually Perturbed Texts for Defending Phishing Attacks
von: Lee, Hanyong, et al.
Veröffentlicht: (2025)
von: Lee, Hanyong, et al.
Veröffentlicht: (2025)
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
EveGuard: Defeating Vibration-based Side-Channel Eavesdropping with Audio Adversarial Perturbations
von: Chang, Jung-Woo, et al.
Veröffentlicht: (2024)
von: Chang, Jung-Woo, et al.
Veröffentlicht: (2024)
Defending Against Sophisticated Poisoning Attacks with RL-based Aggregation in Federated Learning
von: Wang, Yujing, et al.
Veröffentlicht: (2024)
von: Wang, Yujing, et al.
Veröffentlicht: (2024)
EnCAgg: Enhanced Clustering Aggregation for Robust Federated Learning against Dynamic Model Poisoning
von: Zhang, Tianyun, et al.
Veröffentlicht: (2026)
von: Zhang, Tianyun, et al.
Veröffentlicht: (2026)
Privacy Loss of Noise Perturbation via Concentration Analysis of A Product Measure
von: Liu, Shuainan, et al.
Veröffentlicht: (2025)
von: Liu, Shuainan, et al.
Veröffentlicht: (2025)
Retrieval-Confused Generation is a Good Defender for Privacy Violation Attack of Large Language Models
von: Peng, Wanli, et al.
Veröffentlicht: (2025)
von: Peng, Wanli, et al.
Veröffentlicht: (2025)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
Secure Distributed Learning for CAVs: Defending Against Gradient Leakage with Leveled Homomorphic Encryption
von: Najjar, Muhammad Ali, et al.
Veröffentlicht: (2025)
von: Najjar, Muhammad Ali, et al.
Veröffentlicht: (2025)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
von: Teng, Ma, et al.
Veröffentlicht: (2024)
von: Teng, Ma, et al.
Veröffentlicht: (2024)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
von: Yin, Ziyi, et al.
Veröffentlicht: (2025)
von: Yin, Ziyi, et al.
Veröffentlicht: (2025)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
Defend LLMs Through Self-Consciousness
von: Huang, Boshi, et al.
Veröffentlicht: (2025)
von: Huang, Boshi, et al.
Veröffentlicht: (2025)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
von: Guan, Shaowei, et al.
Veröffentlicht: (2025)
von: Guan, Shaowei, et al.
Veröffentlicht: (2025)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
von: Zhao, Yunhan, et al.
Veröffentlicht: (2026)
von: Zhao, Yunhan, et al.
Veröffentlicht: (2026)
Cross-Modal Backdoors in Multimodal Large Language Models
von: Wang, Runhe, et al.
Veröffentlicht: (2026)
von: Wang, Runhe, et al.
Veröffentlicht: (2026)
JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Defending against Jailbreak through Early Exit Generation of Large Language Models
von: Zhao, Chongwen, et al.
Veröffentlicht: (2024) -
KinGuard: Hierarchical Kinship-Aware Fingerprinting to Defend Against Large Language Model Stealing
von: Xu, Zhenhua, et al.
Veröffentlicht: (2026) -
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024) -
Recent Advances in Attack and Defense Approaches of Large Language Models
von: Cui, Jing, et al.
Veröffentlicht: (2024) -
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
von: Zhang, Shenyi, et al.
Veröffentlicht: (2025)