Exploiting Novel GPT-4 APIs
Fuente:
arXiv
Saved in:
| Main Authors: | Pelrine, Kellin, Taufeeque, Mohammad, Zając, Michał, McLean, Euan, Gleave, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Trends in Language Model Robustness
by: Howe, Nikolaus, et al.
Published: (2024)
by: Howe, Nikolaus, et al.
Published: (2024)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026)
by: Struppek, Lukas, et al.
Published: (2026)
Large Language Models are Advanced Anonymizers
by: Staab, Robin, et al.
Published: (2024)
by: Staab, Robin, et al.
Published: (2024)
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
by: Xin, Yuan, et al.
Published: (2025)
by: Xin, Yuan, et al.
Published: (2025)
Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
by: Pai, Aaditya
Published: (2026)
by: Pai, Aaditya
Published: (2026)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Temporal Attack Pattern Detection in Multi-Agent AI Workflows: An Open Framework for Training Trace-Based Security Models
by: Del Rosario, Ron F.
Published: (2025)
by: Del Rosario, Ron F.
Published: (2025)
Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
ConfusionPrompt: Practical Private Inference for Online Large Language Models
by: Mai, Peihua, et al.
Published: (2023)
by: Mai, Peihua, et al.
Published: (2023)
Split-and-Denoise: Protect large language model inference with local differential privacy
by: Mai, Peihua, et al.
Published: (2023)
by: Mai, Peihua, et al.
Published: (2023)
Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
by: Adiletta, Andrew, et al.
Published: (2025)
by: Adiletta, Andrew, et al.
Published: (2025)
AI Safeguards, Generative AI and the Pandora Box: AI Safety Measures to Protect Businesses and Personal Reputation
by: Kumar, Prasanna
Published: (2026)
by: Kumar, Prasanna
Published: (2026)
Efficient LLM Safety Evaluation through Multi-Agent Debate
by: Lin, Dachuan, et al.
Published: (2025)
by: Lin, Dachuan, et al.
Published: (2025)
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
by: Nakamura, Mason, et al.
Published: (2025)
by: Nakamura, Mason, et al.
Published: (2025)
Adversarial Attacks on Large Language Models Using Regularized Relaxation
by: Chacko, Samuel Jacob, et al.
Published: (2024)
by: Chacko, Samuel Jacob, et al.
Published: (2024)
Unlocking the Potential of Large Language Models for Clinical Text Anonymization: A Comparative Study
by: Pissarra, David, et al.
Published: (2024)
by: Pissarra, David, et al.
Published: (2024)
Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design
by: Charles, Richard M., et al.
Published: (2025)
by: Charles, Richard M., et al.
Published: (2025)
Powerful Training-Free Membership Inference Against Autoregressive Language Models
by: Ilić, David, et al.
Published: (2026)
by: Ilić, David, et al.
Published: (2026)
Quantifying Return on Security Controls in LLM Systems
by: Moulton, Richard Helder, et al.
Published: (2025)
by: Moulton, Richard Helder, et al.
Published: (2025)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
by: Shaheer, Safwan, et al.
Published: (2025)
by: Shaheer, Safwan, et al.
Published: (2025)
sudoLLM: On Multi-role Alignment of Language Models
by: Saha, Soumadeep, et al.
Published: (2025)
by: Saha, Soumadeep, et al.
Published: (2025)
UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
by: Geng, Runpeng, et al.
Published: (2025)
by: Geng, Runpeng, et al.
Published: (2025)
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
by: Verma, Apurv, et al.
Published: (2024)
by: Verma, Apurv, et al.
Published: (2024)
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
by: Wang, Xinhai, et al.
Published: (2026)
by: Wang, Xinhai, et al.
Published: (2026)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
by: Miculicich, Lesly, et al.
Published: (2025)
by: Miculicich, Lesly, et al.
Published: (2025)
Safeguarding Efficacy in Large Language Models: Evaluating Resistance to Human-Written and Algorithmic Adversarial Prompts
by: Downey-Webb, Tiarnaigh, et al.
Published: (2025)
by: Downey-Webb, Tiarnaigh, et al.
Published: (2025)
Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
by: Hackett, William, et al.
Published: (2025)
by: Hackett, William, et al.
Published: (2025)
JavelinGuard: Low-Cost Transformer Architectures for LLM Security
by: Datta, Yash, et al.
Published: (2025)
by: Datta, Yash, et al.
Published: (2025)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
by: Palit, Sayon, et al.
Published: (2025)
by: Palit, Sayon, et al.
Published: (2025)
Detecting Prompt Injection Attacks Against Application Using Classifiers
by: Shaheer, Safwan, et al.
Published: (2025)
by: Shaheer, Safwan, et al.
Published: (2025)
SALLIE: Safeguarding Against Latent Language & Image Exploits
by: Azov, Guy, et al.
Published: (2026)
by: Azov, Guy, et al.
Published: (2026)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
Logits of API-Protected LLMs Leak Proprietary Information
by: Finlayson, Matthew, et al.
Published: (2024)
by: Finlayson, Matthew, et al.
Published: (2024)
Shortcuts Arising from Contrast: Effective and Covert Clean-Label Attacks in Prompt-Based Learning
by: Xie, Xiaopeng, et al.
Published: (2024)
by: Xie, Xiaopeng, et al.
Published: (2024)
Train to Defend: First Defense Against Cryptanalytic Neural Network Parameter Extraction Attacks
by: Kurian, Ashley, et al.
Published: (2025)
by: Kurian, Ashley, et al.
Published: (2025)
Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
by: Usman, Rana Muhammad
Published: (2026)
by: Usman, Rana Muhammad
Published: (2026)
Send to which account? Evaluation of an LLM-based Scambaiting System
by: Siadati, Hossein, et al.
Published: (2025)
by: Siadati, Hossein, et al.
Published: (2025)
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
by: Chona, Alankrit, et al.
Published: (2026)
by: Chona, Alankrit, et al.
Published: (2026)
Guardians of the Web: The Evolution and Future of Website Information Security
by: Islam, Md Saiful, et al.
Published: (2025)
by: Islam, Md Saiful, et al.
Published: (2025)
Similar Items
-
Scaling Trends in Language Model Robustness
by: Howe, Nikolaus, et al.
Published: (2024) -
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026) -
Large Language Models are Advanced Anonymizers
by: Staab, Robin, et al.
Published: (2024) -
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
by: Xin, Yuan, et al.
Published: (2025) -
Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
by: Pai, Aaditya
Published: (2026)