From static to adaptive: immune memory-based jailbreak detection for large language models
Fuente:
arXiv
Salvato in:
| Autori principali: | Leng, Jun, Liu, Yu, Zhang, Litian, Hu, Ruihan, Fang, Zhuting, Zhang, Xi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents
di: Shavit, Doron
Pubblicazione: (2026)
di: Shavit, Doron
Pubblicazione: (2026)
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
di: Yi, Xin, et al.
Pubblicazione: (2025)
di: Yi, Xin, et al.
Pubblicazione: (2025)
TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
di: Chen, Xiuyuan, et al.
Pubblicazione: (2025)
di: Chen, Xiuyuan, et al.
Pubblicazione: (2025)
When Reasoning Leaks Membership: Membership Inference Attack on Black-box Large Reasoning Models
di: Hu, Ruihan, et al.
Pubblicazione: (2026)
di: Hu, Ruihan, et al.
Pubblicazione: (2026)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
di: Chen, Zejian, et al.
Pubblicazione: (2026)
di: Chen, Zejian, et al.
Pubblicazione: (2026)
BadEdit: Backdooring large language models by model editing
di: Li, Yanzhou, et al.
Pubblicazione: (2024)
di: Li, Yanzhou, et al.
Pubblicazione: (2024)
Leveraging large language models for SQL behavior-based database intrusion detection
di: Shlezinger, Meital, et al.
Pubblicazione: (2025)
di: Shlezinger, Meital, et al.
Pubblicazione: (2025)
Certified Robust Accuracy of Neural Networks Are Bounded due to Bayes Errors
di: Zhang, Ruihan, et al.
Pubblicazione: (2024)
di: Zhang, Ruihan, et al.
Pubblicazione: (2024)
Feature graph construction with static features for malware detection
di: Zou, Binghui, et al.
Pubblicazione: (2024)
di: Zou, Binghui, et al.
Pubblicazione: (2024)
Vul-LMGNNs: Fusing language models and online-distilled graph neural networks for code vulnerability detection
di: Liu, Ruitong, et al.
Pubblicazione: (2024)
di: Liu, Ruitong, et al.
Pubblicazione: (2024)
PMANet: Malicious URL detection via post-trained language model guided multi-level feature attention network
di: Liu, Ruitong, et al.
Pubblicazione: (2023)
di: Liu, Ruitong, et al.
Pubblicazione: (2023)
MirrorShield: Towards Universal Defense Against Jailbreaks via Entropy-Guided Mirror Crafting
di: Pu, Rui, et al.
Pubblicazione: (2025)
di: Pu, Rui, et al.
Pubblicazione: (2025)
Next-generation cyberattack detection with large language models: anomaly analysis across heterogeneous logs
di: Chagna, Yassine, et al.
Pubblicazione: (2026)
di: Chagna, Yassine, et al.
Pubblicazione: (2026)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
di: Liu, Songyang, et al.
Pubblicazione: (2025)
di: Liu, Songyang, et al.
Pubblicazione: (2025)
Harnessing large-language models to generate private synthetic text
di: Kurakin, Alexey, et al.
Pubblicazione: (2023)
di: Kurakin, Alexey, et al.
Pubblicazione: (2023)
Optimizing watermarks for large language models
di: Wouters, Bram
Pubblicazione: (2023)
di: Wouters, Bram
Pubblicazione: (2023)
S3CDM: A secret-sharing-scheme-based cyberattack detection model and its simulation implementation
di: Chum, Chi Sing, et al.
Pubblicazione: (2026)
di: Chum, Chi Sing, et al.
Pubblicazione: (2026)
Can large language models be privacy preserving and fair medical coders?
di: Dadsetan, Ali, et al.
Pubblicazione: (2024)
di: Dadsetan, Ali, et al.
Pubblicazione: (2024)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
di: Liu, Songyang, et al.
Pubblicazione: (2026)
di: Liu, Songyang, et al.
Pubblicazione: (2026)
Investigating cybersecurity incidents using large language models in latest-generation wireless networks
di: Legashev, Leonid, et al.
Pubblicazione: (2025)
di: Legashev, Leonid, et al.
Pubblicazione: (2025)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
di: Liu, Songyang, et al.
Pubblicazione: (2026)
di: Liu, Songyang, et al.
Pubblicazione: (2026)
A unit-based symbolic execution method for detecting memory corruption vulnerabilities in executable codes
di: Baradaran, Sara, et al.
Pubblicazione: (2022)
di: Baradaran, Sara, et al.
Pubblicazione: (2022)
Hardware-based stack buffer overflow attack detection on RISC-V architectures
di: Chenet, Cristiano Pegoraro, et al.
Pubblicazione: (2024)
di: Chenet, Cristiano Pegoraro, et al.
Pubblicazione: (2024)
An incremental hybrid adaptive network-based IDS in Software Defined Networks to detect stealth attacks
di: Alqahtani, Abdullah H
Pubblicazione: (2024)
di: Alqahtani, Abdullah H
Pubblicazione: (2024)
Time-based GNSS attack detection
di: Spanghero, Marco, et al.
Pubblicazione: (2025)
di: Spanghero, Marco, et al.
Pubblicazione: (2025)
Assessing biomedical knowledge robustness in large language models by query-efficient sampling attacks
di: Xian, R. Patrick, et al.
Pubblicazione: (2024)
di: Xian, R. Patrick, et al.
Pubblicazione: (2024)
"I Don't Use AI for Everything": Exploring Utility, Attitude, and Responsibility of AI-empowered Tools in Software Development
di: Pan, Shidong, et al.
Pubblicazione: (2024)
di: Pan, Shidong, et al.
Pubblicazione: (2024)
From Alerts to Intelligence: A Novel LLM-Aided Framework for Host-based Intrusion Detection
di: Sun, Danyu, et al.
Pubblicazione: (2025)
di: Sun, Danyu, et al.
Pubblicazione: (2025)
Exploring the limits of strong membership inference attacks on large language models
di: Hayes, Jamie, et al.
Pubblicazione: (2025)
di: Hayes, Jamie, et al.
Pubblicazione: (2025)
Deep fused flow and topology features for botnet detection basing on pretrained GCN
di: Xiaoyuan, Meng, et al.
Pubblicazione: (2023)
di: Xiaoyuan, Meng, et al.
Pubblicazione: (2023)
AI-Driven IRM: Transforming insider risk management with adaptive scoring and LLM-based threat detection
di: Koli, Lokesh, et al.
Pubblicazione: (2025)
di: Koli, Lokesh, et al.
Pubblicazione: (2025)
EditMark: Watermarking Large Language Models based on Model Editing
di: Li, Shuai, et al.
Pubblicazione: (2025)
di: Li, Shuai, et al.
Pubblicazione: (2025)
Proof-of-Authorship for Diffusion-based AI Generated Content
di: Lee, De Zhang, et al.
Pubblicazione: (2026)
di: Lee, De Zhang, et al.
Pubblicazione: (2026)
Practically adaptable CPABE based Health-Records sharing framework
di: Imam, Raza, et al.
Pubblicazione: (2024)
di: Imam, Raza, et al.
Pubblicazione: (2024)
Attack and defense techniques in large language models: A survey and new perspectives
di: Liao, Zhiyu, et al.
Pubblicazione: (2025)
di: Liao, Zhiyu, et al.
Pubblicazione: (2025)
Rényi Pufferfish Privacy with Gaussian-based Priors: From Single Gaussian to Mixture Model
di: Yang, Wenjin, et al.
Pubblicazione: (2026)
di: Yang, Wenjin, et al.
Pubblicazione: (2026)
A survey on hardware-based malware detection approaches
di: Chenet, Cristiano Pegoraro, et al.
Pubblicazione: (2023)
di: Chenet, Cristiano Pegoraro, et al.
Pubblicazione: (2023)
Data sharing in the metaverse with key abuse resistance based on decentralized CP-ABE
di: Zhang, Liang, et al.
Pubblicazione: (2024)
di: Zhang, Liang, et al.
Pubblicazione: (2024)
Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Private
di: Wu, Ruihan, et al.
Pubblicazione: (2025)
di: Wu, Ruihan, et al.
Pubblicazione: (2025)
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation
di: Yan, Bingyu, et al.
Pubblicazione: (2026)
di: Yan, Bingyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents
di: Shavit, Doron
Pubblicazione: (2026) -
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
di: Yi, Xin, et al.
Pubblicazione: (2025) -
TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
di: Chen, Xiuyuan, et al.
Pubblicazione: (2025) -
When Reasoning Leaks Membership: Membership Inference Attack on Black-box Large Reasoning Models
di: Hu, Ruihan, et al.
Pubblicazione: (2026) -
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
di: Chen, Zejian, et al.
Pubblicazione: (2026)