Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Inie, Nanna, Stray, Jonathan, Derczynski, Leon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hacc-Man: An Arcade Game for Jailbreaking LLMs
von: Valentim, Matheus, et al.
Veröffentlicht: (2024)
von: Valentim, Matheus, et al.
Veröffentlicht: (2024)
garak: A Framework for Security Probing Large Language Models
von: Derczynski, Leon, et al.
Veröffentlicht: (2024)
von: Derczynski, Leon, et al.
Veröffentlicht: (2024)
Training a General Purpose Automated Red Teaming Model
von: Padmakumar, Aishwarya, et al.
Veröffentlicht: (2026)
von: Padmakumar, Aishwarya, et al.
Veröffentlicht: (2026)
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
What Motivates People to Trust 'AI' Systems?
von: Inie, Nanna
Veröffentlicht: (2024)
von: Inie, Nanna
Veröffentlicht: (2024)
Human-in-the-Loop Generation of Adversarial Texts: A Case Study on Tibetan Script
von: Cao, Xi, et al.
Veröffentlicht: (2024)
von: Cao, Xi, et al.
Veröffentlicht: (2024)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
von: Feng, Yingchaojie, et al.
Veröffentlicht: (2024)
von: Feng, Yingchaojie, et al.
Veröffentlicht: (2024)
Jaco: An Offline Running Privacy-aware Voice Assistant
von: Bermuth, Daniel, et al.
Veröffentlicht: (2022)
von: Bermuth, Daniel, et al.
Veröffentlicht: (2022)
Test Security in Remote Testing Age: Perspectives from Process Data Analytics and AI
von: Hao, Jiangang, et al.
Veröffentlicht: (2024)
von: Hao, Jiangang, et al.
Veröffentlicht: (2024)
Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions
von: Prakash, Vijay, et al.
Veröffentlicht: (2024)
von: Prakash, Vijay, et al.
Veröffentlicht: (2024)
OpenAI's Approach to External Red Teaming for AI Models and Systems
von: Ahmad, Lama, et al.
Veröffentlicht: (2025)
von: Ahmad, Lama, et al.
Veröffentlicht: (2025)
From Theory to Comprehension: A Comparative Study of Differential Privacy and $k$-Anonymity
von: von Voigt, Saskia Nuñez, et al.
Veröffentlicht: (2024)
von: von Voigt, Saskia Nuñez, et al.
Veröffentlicht: (2024)
Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction of Phishing Training Inefficacy Grounded in the NIST Phish Scale
von: Rozema, Andrew T., et al.
Veröffentlicht: (2025)
von: Rozema, Andrew T., et al.
Veröffentlicht: (2025)
Multiverse Privacy Theory for Contextual Risks in Complex User-AI Interactions
von: Gumusel, Ece
Veröffentlicht: (2025)
von: Gumusel, Ece
Veröffentlicht: (2025)
Nudging Users to Change Breached Passwords Using the Protection Motivation Theory
von: Zou, Yixin, et al.
Veröffentlicht: (2024)
von: Zou, Yixin, et al.
Veröffentlicht: (2024)
SoK: Come Together -- Unifying Security, Information Theory, and Cognition for a Mixed Reality Deception Attack Ontology & Analysis Framework
von: Teymourian, Ali, et al.
Veröffentlicht: (2025)
von: Teymourian, Ali, et al.
Veröffentlicht: (2025)
On the Suitability of LLM-Driven Agents for Dark Pattern Audits
von: Sun, Chen, et al.
Veröffentlicht: (2026)
von: Sun, Chen, et al.
Veröffentlicht: (2026)
LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2026)
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2026)
Assessing LLM Response Quality in the Context of Technology-Facilitated Abuse
von: Prakash, Vijay, et al.
Veröffentlicht: (2026)
von: Prakash, Vijay, et al.
Veröffentlicht: (2026)
From Coordinates to Context: An LLM-Bootstrapped Semantic Encoding Framework for Privacy-Preserving Mobile Sensing Stress Recognition
von: Phan, Hoang Khang, et al.
Veröffentlicht: (2025)
von: Phan, Hoang Khang, et al.
Veröffentlicht: (2025)
Emergent misalignment as prompt sensitivity: A research note
von: Wyse, Tim, et al.
Veröffentlicht: (2025)
von: Wyse, Tim, et al.
Veröffentlicht: (2025)
Can LLM-Generated Misinformation Be Detected?
von: Chen, Canyu, et al.
Veröffentlicht: (2023)
von: Chen, Canyu, et al.
Veröffentlicht: (2023)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2023)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2023)
Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis
von: Esposito, Matteo, et al.
Veröffentlicht: (2024)
von: Esposito, Matteo, et al.
Veröffentlicht: (2024)
Usability Study of Security Features in Programmable Logic Controllers
von: Li, Karen, et al.
Veröffentlicht: (2022)
von: Li, Karen, et al.
Veröffentlicht: (2022)
Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active Learning
von: Qiu, Wenjun, et al.
Veröffentlicht: (2024)
von: Qiu, Wenjun, et al.
Veröffentlicht: (2024)
Evaluating the Usability of LLMs in Threat Intelligence Enrichment
von: Srikanth, Sanchana, et al.
Veröffentlicht: (2024)
von: Srikanth, Sanchana, et al.
Veröffentlicht: (2024)
A Survey of Wireless Sensing Security from a Role-Based View: Victim, Weapon, and Shield
von: Geng, Ruixu, et al.
Veröffentlicht: (2024)
von: Geng, Ruixu, et al.
Veröffentlicht: (2024)
Learning from Mistakes: Can LLM Self-Recover after Misalignment?
von: Sorokoletova, Olga E., et al.
Veröffentlicht: (2026)
von: Sorokoletova, Olga E., et al.
Veröffentlicht: (2026)
A Quantitative Study of SMS Phishing Detection
von: Timko, Daniel, et al.
Veröffentlicht: (2023)
von: Timko, Daniel, et al.
Veröffentlicht: (2023)
Towards a Cognitive-Support Tool for Threat Hunters
von: Milani, Alessandra Maciel Paz, et al.
Veröffentlicht: (2026)
von: Milani, Alessandra Maciel Paz, et al.
Veröffentlicht: (2026)
BioShield: A Context-Aware Firewall for Securing Bio-LLMs
von: Das, Protiva, et al.
Veröffentlicht: (2026)
von: Das, Protiva, et al.
Veröffentlicht: (2026)
Grant, Verify, Revoke: A User-Centric Pattern for Blockchain Compliance
von: Khadka, Supriya, et al.
Veröffentlicht: (2026)
von: Khadka, Supriya, et al.
Veröffentlicht: (2026)
Improving Users' Passwords with DPAR: a Data-driven Password Recommendation System
von: Morag, Assaf, et al.
Veröffentlicht: (2024)
von: Morag, Assaf, et al.
Veröffentlicht: (2024)
MORPHEUS: A Multidimensional Framework for Modeling, Measuring, and Mitigating Human Factors in Cybersecurity
von: Desolda, Giuseppe, et al.
Veröffentlicht: (2025)
von: Desolda, Giuseppe, et al.
Veröffentlicht: (2025)
Comparative Simulation of Phishing Attacks on a Critical Information Infrastructure Organization: An Empirical Study
von: Sirawongphatsara, Patsita, et al.
Veröffentlicht: (2024)
von: Sirawongphatsara, Patsita, et al.
Veröffentlicht: (2024)
Actionable Cybersecurity Notifications for Smart Homes: A User Study on the Role of Length and Complexity
von: Jüttner, Victor, et al.
Veröffentlicht: (2025)
von: Jüttner, Victor, et al.
Veröffentlicht: (2025)
Engineering Trust, Creating Vulnerability: A Socio-Technical Analysis of AI Interface Design
von: Kereopa-Yorke, Ben
Veröffentlicht: (2025)
von: Kereopa-Yorke, Ben
Veröffentlicht: (2025)
Hidden-in-Plain-Text: A Benchmark for Social-Web Indirect Prompt Injection in RAG
von: Guo, Haoze, et al.
Veröffentlicht: (2026)
von: Guo, Haoze, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Hacc-Man: An Arcade Game for Jailbreaking LLMs
von: Valentim, Matheus, et al.
Veröffentlicht: (2024) -
garak: A Framework for Security Probing Large Language Models
von: Derczynski, Leon, et al.
Veröffentlicht: (2024) -
Training a General Purpose Automated Red Teaming Model
von: Padmakumar, Aishwarya, et al.
Veröffentlicht: (2026) -
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
von: Chen, Chaoran, et al.
Veröffentlicht: (2025) -
The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)