Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
Fuente:
arXiv
Salvato in:
| Autori principali: | Zou, Andy, Lin, Maxwell, Jones, Eliot, Nowak, Micha, Dziemian, Mateusz, Winter, Nick, Grattan, Alexander, Nathanael, Valent, Croft, Ayla, Davies, Xander, Patel, Jai, Kirk, Robert, Burnikell, Nate, Gal, Yarin, Hendrycks, Dan, Kolter, J. Zico, Fredrikson, Matt |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
di: Krishna, Satyapriya, et al.
Pubblicazione: (2025)
di: Krishna, Satyapriya, et al.
Pubblicazione: (2025)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
di: Feng, Zhili, et al.
Pubblicazione: (2025)
di: Feng, Zhili, et al.
Pubblicazione: (2025)
Evaluating Language Model Reasoning about Confidential Information
di: Sam, Dylan, et al.
Pubblicazione: (2025)
di: Sam, Dylan, et al.
Pubblicazione: (2025)
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
di: Dziemian, Mateusz, et al.
Pubblicazione: (2026)
di: Dziemian, Mateusz, et al.
Pubblicazione: (2026)
Improving Alignment and Robustness with Circuit Breakers
di: Zou, Andy, et al.
Pubblicazione: (2024)
di: Zou, Andy, et al.
Pubblicazione: (2024)
Adversarial Attacks on Robotic Vision Language Action Models
di: Jones, Eliot Krzysztof, et al.
Pubblicazione: (2025)
di: Jones, Eliot Krzysztof, et al.
Pubblicazione: (2025)
Simple Baselines are Competitive with Code Evolution
di: Gideoni, Yonatan, et al.
Pubblicazione: (2026)
di: Gideoni, Yonatan, et al.
Pubblicazione: (2026)
Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
di: Lermen, Simon, et al.
Pubblicazione: (2024)
di: Lermen, Simon, et al.
Pubblicazione: (2024)
Mimetic Initialization of MLPs
di: Trockman, Asher, et al.
Pubblicazione: (2026)
di: Trockman, Asher, et al.
Pubblicazione: (2026)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
di: Akinwande, Victor, et al.
Pubblicazione: (2024)
di: Akinwande, Victor, et al.
Pubblicazione: (2024)
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
di: Davies, Xander, et al.
Pubblicazione: (2025)
di: Davies, Xander, et al.
Pubblicazione: (2025)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
di: Williams, Joshua Nathaniel, et al.
Pubblicazione: (2024)
di: Williams, Joshua Nathaniel, et al.
Pubblicazione: (2024)
Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning
di: Huang, Benhao, et al.
Pubblicazione: (2026)
di: Huang, Benhao, et al.
Pubblicazione: (2026)
Why is SAM Robust to Label Noise?
di: Baek, Christina, et al.
Pubblicazione: (2024)
di: Baek, Christina, et al.
Pubblicazione: (2024)
Safety Pretraining: Toward the Next Generation of Safe AI
di: Maini, Pratyush, et al.
Pubblicazione: (2025)
di: Maini, Pratyush, et al.
Pubblicazione: (2025)
Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
di: Lermen, Simon, et al.
Pubblicazione: (2025)
di: Lermen, Simon, et al.
Pubblicazione: (2025)
Boundary Point Jailbreaking of Black-Box LLMs
di: Davies, Xander, et al.
Pubblicazione: (2026)
di: Davies, Xander, et al.
Pubblicazione: (2026)
Predicting the Performance of Black-box LLMs through Follow-up Queries
di: Sam, Dylan, et al.
Pubblicazione: (2025)
di: Sam, Dylan, et al.
Pubblicazione: (2025)
One-Step Diffusion Distillation via Deep Equilibrium Models
di: Geng, Zhengyang, et al.
Pubblicazione: (2023)
di: Geng, Zhengyang, et al.
Pubblicazione: (2023)
Diffusing Differentiable Representations
di: Savani, Yash, et al.
Pubblicazione: (2024)
di: Savani, Yash, et al.
Pubblicazione: (2024)
Generative Posterior Networks for Approximately Bayesian Epistemic Uncertainty Estimation
di: Roderick, Melrose, et al.
Pubblicazione: (2023)
di: Roderick, Melrose, et al.
Pubblicazione: (2023)
LipNeXt: Scaling up Lipschitz-based Certified Robustness to Billion-parameter Models
di: Hu, Kai, et al.
Pubblicazione: (2026)
di: Hu, Kai, et al.
Pubblicazione: (2026)
Iterative Deployment Improves Planning Skills in LLMs
di: Corrêa, Augusto B., et al.
Pubblicazione: (2025)
di: Corrêa, Augusto B., et al.
Pubblicazione: (2025)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
di: Lin, Justin W., et al.
Pubblicazione: (2025)
di: Lin, Justin W., et al.
Pubblicazione: (2025)
The Mixing method: low-rank coordinate descent for semidefinite programming with diagonal constraints
di: Wang, Po-Wei, et al.
Pubblicazione: (2017)
di: Wang, Po-Wei, et al.
Pubblicazione: (2017)
A Simple and Effective Pruning Approach for Large Language Models
di: Sun, Mingjie, et al.
Pubblicazione: (2023)
di: Sun, Mingjie, et al.
Pubblicazione: (2023)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
di: Goyal, Sachin, et al.
Pubblicazione: (2024)
di: Goyal, Sachin, et al.
Pubblicazione: (2024)
Looking beyond the next token
di: Thankaraj, Abitha, et al.
Pubblicazione: (2025)
di: Thankaraj, Abitha, et al.
Pubblicazione: (2025)
Massive Activations in Large Language Models
di: Sun, Mingjie, et al.
Pubblicazione: (2024)
di: Sun, Mingjie, et al.
Pubblicazione: (2024)
Measuring Five-Nines Reliability: Sample-Efficient LLM Evaluation in Saturated Benchmarks
di: Kim, Eungyeup, et al.
Pubblicazione: (2026)
di: Kim, Eungyeup, et al.
Pubblicazione: (2026)
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
di: O'Brien, Kyle, et al.
Pubblicazione: (2025)
di: O'Brien, Kyle, et al.
Pubblicazione: (2025)
Bernoullicity of some skew products with hyperbolic base and Kochergin flow in the fiber
di: Nowak, Mateusz
Pubblicazione: (2026)
di: Nowak, Mateusz
Pubblicazione: (2026)
Language Models Change Facts Based on the Way You Talk
di: Kearney, Matthew, et al.
Pubblicazione: (2025)
di: Kearney, Matthew, et al.
Pubblicazione: (2025)
Do Multilingual LLMs Think In English?
di: Schut, Lisa, et al.
Pubblicazione: (2025)
di: Schut, Lisa, et al.
Pubblicazione: (2025)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
di: Kossen, Jannik, et al.
Pubblicazione: (2023)
di: Kossen, Jannik, et al.
Pubblicazione: (2023)
The Benefits and Risks of Transductive Approaches for AI Fairness
di: Razzak, Muhammed, et al.
Pubblicazione: (2024)
di: Razzak, Muhammed, et al.
Pubblicazione: (2024)
Finetuning CLIP to Reason about Pairwise Differences
di: Sam, Dylan, et al.
Pubblicazione: (2024)
di: Sam, Dylan, et al.
Pubblicazione: (2024)
Weight Ensembling Improves Reasoning in Language Models
di: Dang, Xingyu, et al.
Pubblicazione: (2025)
di: Dang, Xingyu, et al.
Pubblicazione: (2025)
Understanding Augmentation-based Self-Supervised Representation Learning via RKHS Approximation and Regression
di: Zhai, Runtian, et al.
Pubblicazione: (2023)
di: Zhai, Runtian, et al.
Pubblicazione: (2023)
Documenti analoghi
-
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024) -
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
di: Krishna, Satyapriya, et al.
Pubblicazione: (2025) -
Existing Large Language Model Unlearning Evaluations Are Inconclusive
di: Feng, Zhili, et al.
Pubblicazione: (2025) -
Evaluating Language Model Reasoning about Confidential Information
di: Sam, Dylan, et al.
Pubblicazione: (2025) -
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
di: Dziemian, Mateusz, et al.
Pubblicazione: (2026)