Optimizing AI Agent Attacks With Synthetic Data
Fuente:
arXiv
Saved in:
| Main Authors: | Loughridge, Chloe, Colognese, Paul, Griffin, Avery, Tracy, Tyler, Kutasov, Jon, Benton, Joe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Control Protocols for Untrusted AI Agents
by: Kutasov, Jon, et al.
Published: (2025)
by: Kutasov, Jon, et al.
Published: (2025)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
by: Najt, Elle, et al.
Published: (2026)
by: Najt, Elle, et al.
Published: (2026)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
dafny-annotator: AI-Assisted Verification of Dafny Programs
by: Poesia, Gabriel, et al.
Published: (2024)
by: Poesia, Gabriel, et al.
Published: (2024)
Model Spec Midtraining: Improving How Alignment Training Generalizes
by: Li, Chloe, et al.
Published: (2026)
by: Li, Chloe, et al.
Published: (2026)
Efficiently Aligning Language Models with Online Natural Language Feedback
by: Ye, Christine, et al.
Published: (2026)
by: Ye, Christine, et al.
Published: (2026)
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
by: Schaeffer, Joachim, et al.
Published: (2026)
by: Schaeffer, Joachim, et al.
Published: (2026)
Singular Vectors of Attention Heads Align with Features
by: Franco, Gabriel, et al.
Published: (2026)
by: Franco, Gabriel, et al.
Published: (2026)
BashArena: A Control Setting for Highly Privileged AI Agents
by: Kaufman, Adam, et al.
Published: (2025)
by: Kaufman, Adam, et al.
Published: (2025)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
by: Jotautaitė, Monika, et al.
Published: (2026)
by: Jotautaitė, Monika, et al.
Published: (2026)
Toward a Trustworthy Optimization Modeling Agent via Verifiable Synthetic Data Generation
by: Lima, Vinicius, et al.
Published: (2025)
by: Lima, Vinicius, et al.
Published: (2025)
Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
by: Folkerts, Linus, et al.
Published: (2026)
by: Folkerts, Linus, et al.
Published: (2026)
Generative AI in clinical practice: novel qualitative evidence of risk and responsible use of Google's NotebookLM
by: Reuter, Max, et al.
Published: (2025)
by: Reuter, Max, et al.
Published: (2025)
When Privacy Isn't Synthetic: Hidden Data Leakage in Generative AI Models
by: Mustaqim, S. M., et al.
Published: (2025)
by: Mustaqim, S. M., et al.
Published: (2025)
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
by: Turpin, Miles, et al.
Published: (2025)
by: Turpin, Miles, et al.
Published: (2025)
Removing Sandbagging in LLMs by Training with Weak Supervision
by: Ryd, Emil, et al.
Published: (2026)
by: Ryd, Emil, et al.
Published: (2026)
AgentSGEN: Multi-Agent LLM in the Loop for Semantic Collaboration and GENeration of Synthetic Data
by: Xuan, Vu Dinh, et al.
Published: (2025)
by: Xuan, Vu Dinh, et al.
Published: (2025)
Combining Cost-Constrained Runtime Monitors for AI Safety
by: Hua, Tim Tian, et al.
Published: (2025)
by: Hua, Tim Tian, et al.
Published: (2025)
Superplatforms Have to Attack AI Agents
by: Lin, Jianghao, et al.
Published: (2025)
by: Lin, Jianghao, et al.
Published: (2025)
Securing AI Agents Against Prompt Injection Attacks
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
Emerging Cyber Attack Risks of Medical AI Agents
by: Qiu, Jianing, et al.
Published: (2025)
by: Qiu, Jianing, et al.
Published: (2025)
Automatically Attacking Software Reverse Engineering AI Agents
by: Crawford, Brian, et al.
Published: (2026)
by: Crawford, Brian, et al.
Published: (2026)
From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent
by: Wang, Yuhang, et al.
Published: (2026)
by: Wang, Yuhang, et al.
Published: (2026)
CI-Bench: Benchmarking Contextual Integrity of AI Assistants on Synthetic Data
by: Cheng, Zhao, et al.
Published: (2024)
by: Cheng, Zhao, et al.
Published: (2024)
ATAG: AI-Agent Application Threat Assessment with Attack Graphs
by: Gandhi, Parth Atulbhai, et al.
Published: (2025)
by: Gandhi, Parth Atulbhai, et al.
Published: (2025)
Repurposing Synthetic Data for Fine-grained Search Agent Supervision
by: Zhao, Yida, et al.
Published: (2025)
by: Zhao, Yida, et al.
Published: (2025)
Linguistic and Argument Diversity in Synthetic Data for Function-Calling Agents
by: Greenstein, Dan, et al.
Published: (2026)
by: Greenstein, Dan, et al.
Published: (2026)
Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025
by: Huang, Ruanqianqian, et al.
Published: (2025)
by: Huang, Ruanqianqian, et al.
Published: (2025)
TO-Agents: A Multi-Agent AI Pipeline for Preference-Guided Topology Optimization
by: Stewart, Isabella A., et al.
Published: (2026)
by: Stewart, Isabella A., et al.
Published: (2026)
Natural Emergent Misalignment from Reward Hacking in Production RL
by: MacDiarmid, Monte, et al.
Published: (2025)
by: MacDiarmid, Monte, et al.
Published: (2025)
Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
by: Li, Dawei, et al.
Published: (2025)
by: Li, Dawei, et al.
Published: (2025)
Polysemanticity and Capacity in Neural Networks
by: Scherlis, Adam, et al.
Published: (2022)
by: Scherlis, Adam, et al.
Published: (2022)
Synthetic Data for Robust AI Model Development in Regulated Enterprises
by: Godbole, Aditi
Published: (2025)
by: Godbole, Aditi
Published: (2025)
Template-as-Ontology: Configurable Synthetic Data Infrastructure for Cross-Domain Manufacturing AI Validation
by: Chethan, Grama
Published: (2026)
by: Chethan, Grama
Published: (2026)
Synthetic Trust Attacks: Modeling How Generative AI Manipulates Human Decisions in Social Engineering Fraud
by: Ashraf, Muhammad Tahir
Published: (2026)
by: Ashraf, Muhammad Tahir
Published: (2026)
SymptomWise: A Deterministic Reasoning Layer for Reliable and Efficient AI Systems
by: Henry, Isaac, et al.
Published: (2026)
by: Henry, Isaac, et al.
Published: (2026)
Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks
by: Isbarov, Jafar, et al.
Published: (2026)
by: Isbarov, Jafar, et al.
Published: (2026)
Internal Deployment Gaps in AI Regulation
by: Kwon, Joe, et al.
Published: (2026)
by: Kwon, Joe, et al.
Published: (2026)
Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation
by: Fan, Zhiting, et al.
Published: (2026)
by: Fan, Zhiting, et al.
Published: (2026)
Kinematic Tokenization: Optimization-Based Continuous-Time Tokens for Learnable Decision Policies in Noisy Time Series
by: Kearney, Griffin
Published: (2026)
by: Kearney, Griffin
Published: (2026)
Similar Items
-
Evaluating Control Protocols for Untrusted AI Agents
by: Kutasov, Jon, et al.
Published: (2025) -
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
by: Najt, Elle, et al.
Published: (2026) -
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025) -
dafny-annotator: AI-Assisted Verification of Dafny Programs
by: Poesia, Gabriel, et al.
Published: (2024) -
Model Spec Midtraining: Improving How Alignment Training Generalizes
by: Li, Chloe, et al.
Published: (2026)