Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Shavit, Doron |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From static to adaptive: immune memory-based jailbreak detection for large language models
von: Leng, Jun, et al.
Veröffentlicht: (2025)
von: Leng, Jun, et al.
Veröffentlicht: (2025)
Next-generation cyberattack detection with large language models: anomaly analysis across heterogeneous logs
von: Chagna, Yassine, et al.
Veröffentlicht: (2026)
von: Chagna, Yassine, et al.
Veröffentlicht: (2026)
An explainable Recursive Feature Elimination to detect Advanced Persistent Threats using Random Forest classifier
von: Mutalib, Noor Hazlina Abdul, et al.
Veröffentlicht: (2025)
von: Mutalib, Noor Hazlina Abdul, et al.
Veröffentlicht: (2025)
Attacks on the neural network and defense methods
von: Korenev, A., et al.
Veröffentlicht: (2024)
von: Korenev, A., et al.
Veröffentlicht: (2024)
Attack and defense techniques in large language models: A survey and new perspectives
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
von: Chen, Xiuyuan, et al.
Veröffentlicht: (2025)
von: Chen, Xiuyuan, et al.
Veröffentlicht: (2025)
BadEdit: Backdooring large language models by model editing
von: Li, Yanzhou, et al.
Veröffentlicht: (2024)
von: Li, Yanzhou, et al.
Veröffentlicht: (2024)
Enhancing Adversarial Resistance in LLMs with Recursion
von: Li, Bryan, et al.
Veröffentlicht: (2024)
von: Li, Bryan, et al.
Veröffentlicht: (2024)
Network evasion detection with Bi-LSTM model
von: Chen, Kehua, et al.
Veröffentlicht: (2025)
von: Chen, Kehua, et al.
Veröffentlicht: (2025)
Evaluation empirique de la sécurisation et de l'alignement de ChatGPT et Gemini: analyse comparative des vulnérabilités par expérimentations de jailbreaks
von: Nouailles, Rafaël
Veröffentlicht: (2025)
von: Nouailles, Rafaël
Veröffentlicht: (2025)
Backdoor defense, learnability and obfuscation
von: Christiano, Paul, et al.
Veröffentlicht: (2024)
von: Christiano, Paul, et al.
Veröffentlicht: (2024)
Investigating cybersecurity incidents using large language models in latest-generation wireless networks
von: Legashev, Leonid, et al.
Veröffentlicht: (2025)
von: Legashev, Leonid, et al.
Veröffentlicht: (2025)
The importance of the clustering model to detect new types of intrusion in data traffic
von: Abd, Noor Saud, et al.
Veröffentlicht: (2024)
von: Abd, Noor Saud, et al.
Veröffentlicht: (2024)
FL-CLEANER: byzantine and backdoor defense by CLustering Errors of Activation maps in Non-iid fedErated leaRning
von: Ghali, Mehdi Ben, et al.
Veröffentlicht: (2025)
von: Ghali, Mehdi Ben, et al.
Veröffentlicht: (2025)
MURMUR: Using cross-user chatter to break collaborative language agents in groups
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
RECUR: Resource Exhaustion Attack via Recursive-Entropy Guided Counterfactual Utilization and Reflection
von: Wang, Ziwei, et al.
Veröffentlicht: (2026)
von: Wang, Ziwei, et al.
Veröffentlicht: (2026)
Optimized detection of cyber-attacks on IoT networks via hybrid deep learning models
von: Bensaoud, Ahmed, et al.
Veröffentlicht: (2025)
von: Bensaoud, Ahmed, et al.
Veröffentlicht: (2025)
Research and application of artificial intelligence based webshell detection model: A literature review
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
Optimizing watermarks for large language models
von: Wouters, Bram
Veröffentlicht: (2023)
von: Wouters, Bram
Veröffentlicht: (2023)
AI Propaganda factories with language models
von: Olejnik, Lukasz
Veröffentlicht: (2025)
von: Olejnik, Lukasz
Veröffentlicht: (2025)
DAIRE: A lightweight AI model for real-time detection of Controller Area Network attacks in the Internet of Vehicles
von: Alam, Shahid, et al.
Veröffentlicht: (2026)
von: Alam, Shahid, et al.
Veröffentlicht: (2026)
Security awareness in LLM agents: the NDAI zone case
von: Bottazzi, Enrico, et al.
Veröffentlicht: (2026)
von: Bottazzi, Enrico, et al.
Veröffentlicht: (2026)
Prompt Leakage effect and defense strategies for multi-turn LLM interactions
von: Agarwal, Divyansh, et al.
Veröffentlicht: (2024)
von: Agarwal, Divyansh, et al.
Veröffentlicht: (2024)
AI Kill Switch for malicious web-based LLM agent
von: Lee, Sechan, et al.
Veröffentlicht: (2025)
von: Lee, Sechan, et al.
Veröffentlicht: (2025)
Context manipulation attacks : Web agents are susceptible to corrupted memory
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
Multi-agent Reinforcement Learning-based Network Intrusion Detection System
von: Tellache, Amine, et al.
Veröffentlicht: (2024)
von: Tellache, Amine, et al.
Veröffentlicht: (2024)
From surveillance to signalling: escalation channels as environmental controls for agentic AI
von: Gomez, Francesca
Veröffentlicht: (2025)
von: Gomez, Francesca
Veröffentlicht: (2025)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
LlamaFirewall: An open source guardrail system for building secure AI agents
von: Chennabasappa, Sahana, et al.
Veröffentlicht: (2025)
von: Chennabasappa, Sahana, et al.
Veröffentlicht: (2025)
Towards the generation of hierarchical attack models from cybersecurity vulnerabilities using language models
von: Sowka, Kacper, et al.
Veröffentlicht: (2024)
von: Sowka, Kacper, et al.
Veröffentlicht: (2024)
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
von: Yi, Xin, et al.
Veröffentlicht: (2025)
von: Yi, Xin, et al.
Veröffentlicht: (2025)
Forbidden Facts: An Investigation of Competing Objectives in Llama-2
von: Wang, Tony T., et al.
Veröffentlicht: (2023)
von: Wang, Tony T., et al.
Veröffentlicht: (2023)
Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection
von: Halder, Subho, et al.
Veröffentlicht: (2026)
von: Halder, Subho, et al.
Veröffentlicht: (2026)
Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems
von: Ko, Ronny, et al.
Veröffentlicht: (2025)
von: Ko, Ronny, et al.
Veröffentlicht: (2025)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
ActDroid: An active learning framework for Android malware detection
von: Muzaffar, Ali, et al.
Veröffentlicht: (2024)
von: Muzaffar, Ali, et al.
Veröffentlicht: (2024)
Exploring the limits of strong membership inference attacks on large language models
von: Hayes, Jamie, et al.
Veröffentlicht: (2025)
von: Hayes, Jamie, et al.
Veröffentlicht: (2025)
Anomaly detection in network flows using unsupervised online machine learning
von: Miguel-Diez, Alberto, et al.
Veröffentlicht: (2025)
von: Miguel-Diez, Alberto, et al.
Veröffentlicht: (2025)
COPS: A Compact On-device Pipeline for real-time Smishing detection
von: S, Harichandana B S, et al.
Veröffentlicht: (2024)
von: S, Harichandana B S, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
From static to adaptive: immune memory-based jailbreak detection for large language models
von: Leng, Jun, et al.
Veröffentlicht: (2025) -
Next-generation cyberattack detection with large language models: anomaly analysis across heterogeneous logs
von: Chagna, Yassine, et al.
Veröffentlicht: (2026) -
An explainable Recursive Feature Elimination to detect Advanced Persistent Threats using Random Forest classifier
von: Mutalib, Noor Hazlina Abdul, et al.
Veröffentlicht: (2025) -
Attacks on the neural network and defense methods
von: Korenev, A., et al.
Veröffentlicht: (2024) -
Attack and defense techniques in large language models: A survey and new perspectives
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)