DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
Fuente:
arXiv
Salvato in:
| Autori principali: | Boisvert, Leo, Bansal, Mihir, Evuru, Chandra Kiran Reddy, Huang, Gabriel, Puri, Abhay, Bose, Avinandan, Fazel, Maryam, Cappart, Quentin, Stanley, Jason, Lacoste, Alexandre, Drouin, Alexandre, Dvijotham, Krishnamurthy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
di: Boisvert, Léo, et al.
Pubblicazione: (2025)
di: Boisvert, Léo, et al.
Pubblicazione: (2025)
Keeping up with dynamic attackers: Certifying robustness to adaptive online data poisoning
di: Bose, Avinandan, et al.
Pubblicazione: (2025)
di: Bose, Avinandan, et al.
Pubblicazione: (2025)
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
di: Boisvert, Léo, et al.
Pubblicazione: (2024)
di: Boisvert, Léo, et al.
Pubblicazione: (2024)
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
di: Drouin, Alexandre, et al.
Pubblicazione: (2024)
di: Drouin, Alexandre, et al.
Pubblicazione: (2024)
Towards a Generic Representation of Combinatorial Problems for Learning-Based Approaches
di: Boisvert, Léo, et al.
Pubblicazione: (2024)
di: Boisvert, Léo, et al.
Pubblicazione: (2024)
Offline Multi-task Transfer RL with Representational Penalization
di: Bose, Avinandan, et al.
Pubblicazione: (2024)
di: Bose, Avinandan, et al.
Pubblicazione: (2024)
Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?
di: Bhagwatkar, Rishika, et al.
Pubblicazione: (2025)
di: Bhagwatkar, Rishika, et al.
Pubblicazione: (2025)
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
di: Bose, Avinandan, et al.
Pubblicazione: (2024)
di: Bose, Avinandan, et al.
Pubblicazione: (2024)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
di: Kazdan, Joshua, et al.
Pubblicazione: (2025)
di: Kazdan, Joshua, et al.
Pubblicazione: (2025)
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
di: Bose, Avinandan, et al.
Pubblicazione: (2025)
di: Bose, Avinandan, et al.
Pubblicazione: (2025)
LitLLM: A Toolkit for Scientific Literature Review
di: Agarwal, Shubham, et al.
Pubblicazione: (2024)
di: Agarwal, Shubham, et al.
Pubblicazione: (2024)
LitLLMs, LLMs for Literature Review: Are we there yet?
di: Agarwal, Shubham, et al.
Pubblicazione: (2024)
di: Agarwal, Shubham, et al.
Pubblicazione: (2024)
Initializing Services in Interactive ML Systems for Diverse Users
di: Bose, Avinandan, et al.
Pubblicazione: (2023)
di: Bose, Avinandan, et al.
Pubblicazione: (2023)
PrefDisco: Benchmarking Proactive Personalized Reasoning
di: Li, Shuyue Stella, et al.
Pubblicazione: (2025)
di: Li, Shuyue Stella, et al.
Pubblicazione: (2025)
The BrowserGym Ecosystem for Web Agent Research
di: De Chezelles, Thibault Le Sellier, et al.
Pubblicazione: (2024)
di: De Chezelles, Thibault Le Sellier, et al.
Pubblicazione: (2024)
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
di: Sahu, Gaurav, et al.
Pubblicazione: (2024)
di: Sahu, Gaurav, et al.
Pubblicazione: (2024)
Achieving the Tightest Relaxation of Sigmoids for Formal Verification
di: Chevalier, Samuel, et al.
Pubblicazione: (2024)
di: Chevalier, Samuel, et al.
Pubblicazione: (2024)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
di: Nejadgholi, Isar, et al.
Pubblicazione: (2025)
di: Nejadgholi, Isar, et al.
Pubblicazione: (2025)
RECAP: Retrieval-Augmented Audio Captioning
di: Ghosh, Sreyan, et al.
Pubblicazione: (2023)
di: Ghosh, Sreyan, et al.
Pubblicazione: (2023)
Norm-Bounded Low-Rank Adaptation
di: Wang, Ruigang, et al.
Pubblicazione: (2025)
di: Wang, Ruigang, et al.
Pubblicazione: (2025)
Monotone, Bi-Lipschitz, and Polyak-Lojasiewicz Networks
di: Wang, Ruigang, et al.
Pubblicazione: (2024)
di: Wang, Ruigang, et al.
Pubblicazione: (2024)
Is Behavioral Economics Doomed?
di: K. Levine, David
Pubblicazione: (2018)
di: K. Levine, David
Pubblicazione: (2018)
Cold-Start Personalization via Training-Free Priors from Structured World Models
di: Bose, Avinandan, et al.
Pubblicazione: (2026)
di: Bose, Avinandan, et al.
Pubblicazione: (2026)
CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP
di: Evuru, Chandra Kiran Reddy, et al.
Pubblicazione: (2024)
di: Evuru, Chandra Kiran Reddy, et al.
Pubblicazione: (2024)
Global Rewards in Multi-Agent Deep Reinforcement Learning for Autonomous Mobility on Demand Systems
di: Hoppe, Heiko, et al.
Pubblicazione: (2023)
di: Hoppe, Heiko, et al.
Pubblicazione: (2023)
Deep Learning for Data-Driven Districting-and-Routing
di: Ferraz, Arthur, et al.
Pubblicazione: (2024)
di: Ferraz, Arthur, et al.
Pubblicazione: (2024)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
di: Miculicich, Lesly, et al.
Pubblicazione: (2025)
di: Miculicich, Lesly, et al.
Pubblicazione: (2025)
One Color Makes All the Difference in the Tractability of Partial Coloring in Semi-Streaming
di: Das, Avinandan
Pubblicazione: (2026)
di: Das, Avinandan
Pubblicazione: (2026)
Behavioural Cloning in VizDoom
di: Spick, Ryan, et al.
Pubblicazione: (2024)
di: Spick, Ryan, et al.
Pubblicazione: (2024)
The Case Against Climate Doom
di: Jakob, Michael
Pubblicazione: (2025)
di: Jakob, Michael
Pubblicazione: (2025)
Is the System of Nation‐States Doomed?
di: Branko Milanovic
Pubblicazione: (2026)
di: Branko Milanovic
Pubblicazione: (2026)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents
di: Kerboua, Imene, et al.
Pubblicazione: (2025)
di: Kerboua, Imene, et al.
Pubblicazione: (2025)
MARCO: A Memory-Augmented Reinforcement Framework for Combinatorial Optimization
di: Garmendia, Andoni I., et al.
Pubblicazione: (2024)
di: Garmendia, Andoni I., et al.
Pubblicazione: (2024)
An Evaluation of "Choice" as a Selection Tool in the Field of Western History.
di: Boisvert, Marianne
Pubblicazione: (1977)
di: Boisvert, Marianne
Pubblicazione: (1977)
Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems
di: Madhusudhan, Nishanth, et al.
Pubblicazione: (2026)
di: Madhusudhan, Nishanth, et al.
Pubblicazione: (2026)
CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution
di: Kim, Minbeom, et al.
Pubblicazione: (2026)
di: Kim, Minbeom, et al.
Pubblicazione: (2026)
JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
di: Nekoei, Hadi, et al.
Pubblicazione: (2025)
di: Nekoei, Hadi, et al.
Pubblicazione: (2025)
ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations
di: Ghosh, Sreyan, et al.
Pubblicazione: (2023)
di: Ghosh, Sreyan, et al.
Pubblicazione: (2023)
Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
di: Boisvert, Léo, et al.
Pubblicazione: (2025) -
Keeping up with dynamic attackers: Certifying robustness to adaptive online data poisoning
di: Bose, Avinandan, et al.
Pubblicazione: (2025) -
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
di: Boisvert, Léo, et al.
Pubblicazione: (2024) -
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
di: Drouin, Alexandre, et al.
Pubblicazione: (2024) -
Towards a Generic Representation of Combinatorial Problems for Learning-Based Approaches
di: Boisvert, Léo, et al.
Pubblicazione: (2024)