In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Durner, Nils |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
gpt-oss-120b & gpt-oss-20b Model Card
von: OpenAI, et al.
Veröffentlicht: (2025)
von: OpenAI, et al.
Veröffentlicht: (2025)
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
von: Jin, Xisen, et al.
Veröffentlicht: (2026)
von: Jin, Xisen, et al.
Veröffentlicht: (2026)
A Large-Scale Empirical Analysis of Custom GPTs' Vulnerabilities in the OpenAI Ecosystem
von: Ogundoyin, Sunday Oyinlola, et al.
Veröffentlicht: (2025)
von: Ogundoyin, Sunday Oyinlola, et al.
Veröffentlicht: (2025)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
von: Iqbal, Umar, et al.
Veröffentlicht: (2023)
von: Iqbal, Umar, et al.
Veröffentlicht: (2023)
Privacy and Security Threat for OpenAI GPTs
von: Wenying, Wei, et al.
Veröffentlicht: (2025)
von: Wenying, Wei, et al.
Veröffentlicht: (2025)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
OpenAI's Approach to External Red Teaming for AI Models and Systems
von: Ahmad, Lama, et al.
Veröffentlicht: (2025)
von: Ahmad, Lama, et al.
Veröffentlicht: (2025)
OneShield -- the Next Generation of LLM Guardrails
von: DeLuca, Chad, et al.
Veröffentlicht: (2025)
von: DeLuca, Chad, et al.
Veröffentlicht: (2025)
NeuroFilter: Privacy Guardrails for Conversational LLM Agents
von: Das, Saswat, et al.
Veröffentlicht: (2026)
von: Das, Saswat, et al.
Veröffentlicht: (2026)
SGuard-v1: Safety Guardrail for Large Language Models
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
A Comparative Evaluation of AI Agent Security Guardrails
von: Li, Qi, et al.
Veröffentlicht: (2026)
von: Li, Qi, et al.
Veröffentlicht: (2026)
AVISE: Framework for Evaluating the Security of AI Systems
von: Lempinen, Mikko, et al.
Veröffentlicht: (2026)
von: Lempinen, Mikko, et al.
Veröffentlicht: (2026)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models
von: Owiredu-Ashley, Harry
Veröffentlicht: (2026)
von: Owiredu-Ashley, Harry
Veröffentlicht: (2026)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
von: Vega, Jason, et al.
Veröffentlicht: (2023)
von: Vega, Jason, et al.
Veröffentlicht: (2023)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
von: Wang, Libo
Veröffentlicht: (2024)
von: Wang, Libo
Veröffentlicht: (2024)
Current state of LLM Risks and AI Guardrails
von: Ayyamperumal, Suriya Ganesh, et al.
Veröffentlicht: (2024)
von: Ayyamperumal, Suriya Ganesh, et al.
Veröffentlicht: (2024)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
Enhancing Guardrails for Safe and Secure Healthcare AI
von: Gangavarapu, Ananya
Veröffentlicht: (2024)
von: Gangavarapu, Ananya
Veröffentlicht: (2024)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
RedacBench: Can AI Erase Your Secrets?
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
Analysis and prevention of AI-based phishing email attacks
von: Eze, Chibuike Samuel, et al.
Veröffentlicht: (2024)
von: Eze, Chibuike Samuel, et al.
Veröffentlicht: (2024)
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
von: Yuan, Zhuowen, et al.
Veröffentlicht: (2024)
von: Yuan, Zhuowen, et al.
Veröffentlicht: (2024)
VelLMes: A high-interaction AI-based deception framework
von: Sladić, Muris, et al.
Veröffentlicht: (2025)
von: Sladić, Muris, et al.
Veröffentlicht: (2025)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
von: Teja, Lekkala Sai, et al.
Veröffentlicht: (2025)
von: Teja, Lekkala Sai, et al.
Veröffentlicht: (2025)
Code Vulnerability Detection Across Different Programming Languages with AI Models
von: Humran, Hael Abdulhakim Ali, et al.
Veröffentlicht: (2025)
von: Humran, Hael Abdulhakim Ali, et al.
Veröffentlicht: (2025)
What Was Your Prompt? A Remote Keylogging Attack on AI Assistants
von: Weiss, Roy, et al.
Veröffentlicht: (2024)
von: Weiss, Roy, et al.
Veröffentlicht: (2024)
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
von: Schnabl, Christoph, et al.
Veröffentlicht: (2025)
von: Schnabl, Christoph, et al.
Veröffentlicht: (2025)
Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers
von: Bertollo, Giacomo, et al.
Veröffentlicht: (2025)
von: Bertollo, Giacomo, et al.
Veröffentlicht: (2025)
Using Hallucinations to Bypass GPT4's Filter
von: Lemkin, Benjamin
Veröffentlicht: (2024)
von: Lemkin, Benjamin
Veröffentlicht: (2024)
Guarding Your Conversations: Privacy Gatekeepers for Secure Interactions with Cloud-Based AI Models
von: Uzor, GodsGift, et al.
Veröffentlicht: (2025)
von: Uzor, GodsGift, et al.
Veröffentlicht: (2025)
AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation
von: Roy, Joyjit, et al.
Veröffentlicht: (2026)
von: Roy, Joyjit, et al.
Veröffentlicht: (2026)
Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
RvB: Automating AI System Hardening via Iterative Red-Blue Games
von: Huang, Lige, et al.
Veröffentlicht: (2026)
von: Huang, Lige, et al.
Veröffentlicht: (2026)
Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms
von: Azarafrooz, Ari
Veröffentlicht: (2026)
von: Azarafrooz, Ari
Veröffentlicht: (2026)
OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models
von: Wang, Thomas, et al.
Veröffentlicht: (2025)
von: Wang, Thomas, et al.
Veröffentlicht: (2025)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
Cognitive Control Architecture (CCA): A Lifecycle Supervision Framework for Robustly Aligned AI Agents
von: Liang, Zhibo, et al.
Veröffentlicht: (2025)
von: Liang, Zhibo, et al.
Veröffentlicht: (2025)
PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
von: Munoz, Gary D. Lopez, et al.
Veröffentlicht: (2024)
von: Munoz, Gary D. Lopez, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
gpt-oss-120b & gpt-oss-20b Model Card
von: OpenAI, et al.
Veröffentlicht: (2025) -
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
von: Jin, Xisen, et al.
Veröffentlicht: (2026) -
A Large-Scale Empirical Analysis of Custom GPTs' Vulnerabilities in the OpenAI Ecosystem
von: Ogundoyin, Sunday Oyinlola, et al.
Veröffentlicht: (2025) -
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
von: Iqbal, Umar, et al.
Veröffentlicht: (2023) -
Privacy and Security Threat for OpenAI GPTs
von: Wenying, Wei, et al.
Veröffentlicht: (2025)