Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nöther, Jonathan, Singla, Adish, Radanovic, Goran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
von: Nöther, Jonathan, et al.
Veröffentlicht: (2026)
von: Nöther, Jonathan, et al.
Veröffentlicht: (2026)
Policy Teaching via Data Poisoning in Learning from Human Preferences
von: Nika, Andi, et al.
Veröffentlicht: (2025)
von: Nika, Andi, et al.
Veröffentlicht: (2025)
Corruption-Robust Offline Two-Player Zero-Sum Markov Games
von: Nika, Andi, et al.
Veröffentlicht: (2024)
von: Nika, Andi, et al.
Veröffentlicht: (2024)
Corruption Robust Offline Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
Learning Embeddings for Sequential Tasks Using Population of Agents
von: Mahajan, Mridul, et al.
Veröffentlicht: (2023)
von: Mahajan, Mridul, et al.
Veröffentlicht: (2023)
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
von: Nika, Andi, et al.
Veröffentlicht: (2026)
von: Nika, Andi, et al.
Veröffentlicht: (2026)
AgenticRed: Evolving Agentic Systems for Red-Teaming
von: Yuan, Jiayi, et al.
Veröffentlicht: (2026)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2026)
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
von: Nika, Andi, et al.
Veröffentlicht: (2024)
von: Nika, Andi, et al.
Veröffentlicht: (2024)
On Corruption-Robustness in Performative Reinforcement Learning
von: Pollatos, Vasilis, et al.
Veröffentlicht: (2025)
von: Pollatos, Vasilis, et al.
Veröffentlicht: (2025)
Hints-In-Browser: Benchmarking Language Models for Programming Feedback Generation
von: Kotalwar, Nachiket, et al.
Veröffentlicht: (2024)
von: Kotalwar, Nachiket, et al.
Veröffentlicht: (2024)
Distributionally Robust Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2025)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2025)
Reward Design for Justifiable Sequential Decision-Making
von: Sukovic, Aleksa, et al.
Veröffentlicht: (2024)
von: Sukovic, Aleksa, et al.
Veröffentlicht: (2024)
Performative Reinforcement Learning with Linear Markov Decision Process
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2025)
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2025)
Informativeness of Reward Functions in Reinforcement Learning
von: Devidze, Rati, et al.
Veröffentlicht: (2024)
von: Devidze, Rati, et al.
Veröffentlicht: (2024)
Proximal Curriculum with Task Correlations for Deep Reinforcement Learning
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2024)
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2024)
Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
von: Sasnauskas, Paulius, et al.
Veröffentlicht: (2025)
von: Sasnauskas, Paulius, et al.
Veröffentlicht: (2025)
Towards Generalizable Agents in Text-Based Educational Environments: A Study of Integrating RL with LLMs
von: Radmehr, Bahar, et al.
Veröffentlicht: (2024)
von: Radmehr, Bahar, et al.
Veröffentlicht: (2024)
Sparse Offline Reinforcement Learning with Corruption Robustness
von: Tran, Nam Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Nam Phuong, et al.
Veröffentlicht: (2025)
Performative Reinforcement Learning in Gradually Shifting Environments
von: Rank, Ben, et al.
Veröffentlicht: (2024)
von: Rank, Ben, et al.
Veröffentlicht: (2024)
Independent Learning in Performative Markov Potential Games
von: Sahitaj, Rilind, et al.
Veröffentlicht: (2025)
von: Sahitaj, Rilind, et al.
Veröffentlicht: (2025)
The Surprising Harmfulness of Benign Overfitting for Adversarial Robustness
von: Hao, Yifan, et al.
Veröffentlicht: (2024)
von: Hao, Yifan, et al.
Veröffentlicht: (2024)
Neural Task Synthesis for Visual Programming
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2023)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2023)
Formal Models of Active Learning from Contrastive Examples
von: Mansouri, Farnam, et al.
Veröffentlicht: (2025)
von: Mansouri, Farnam, et al.
Veröffentlicht: (2025)
Inference-Time Personalized Alignment with a Few User Preference Queries
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
Stochastic Principal-Agent Problems: Efficient Computation and Learning
von: Gan, Jiarui, et al.
Veröffentlicht: (2023)
von: Gan, Jiarui, et al.
Veröffentlicht: (2023)
Adversarially Robust Detection of Harmful Online Content: A Computational Design Science Approach
von: Chai, Yidong, et al.
Veröffentlicht: (2025)
von: Chai, Yidong, et al.
Veröffentlicht: (2025)
Learning Half-Spaces from Perturbed Contrastive Examples
von: Ravari, Aryan Alavi Razavi, et al.
Veröffentlicht: (2026)
von: Ravari, Aryan Alavi Razavi, et al.
Veröffentlicht: (2026)
Reinforcement Learning for Durable Algorithmic Recourse
von: Ceccon, Marina, et al.
Veröffentlicht: (2025)
von: Ceccon, Marina, et al.
Veröffentlicht: (2025)
Optimal Decision Making Under Strategic Behavior
von: Tsirtsis, Stratis, et al.
Veröffentlicht: (2019)
von: Tsirtsis, Stratis, et al.
Veröffentlicht: (2019)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
Robustness of Agentic AI Systems via Adversarially-Aligned Jacobian Regularization
von: Mumcu, Furkan, et al.
Veröffentlicht: (2026)
von: Mumcu, Furkan, et al.
Veröffentlicht: (2026)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
On the Benefits of Inducing Local Lipschitzness for Robust Generative Adversarial Imitation Learning
von: Memarian, Farzan, et al.
Veröffentlicht: (2021)
von: Memarian, Farzan, et al.
Veröffentlicht: (2021)
Benchmarking Generative Models on Computational Thinking Tests in Elementary Visual Programming
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2024)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2024)
Whispers in the Machine: Confidentiality in Agentic Systems
von: Evertz, Jonathan, et al.
Veröffentlicht: (2024)
von: Evertz, Jonathan, et al.
Veröffentlicht: (2024)
Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
Adversarial Pruning: A Survey and Benchmark of Pruning Methods for Adversarial Robustness
von: Piras, Giorgio, et al.
Veröffentlicht: (2024)
von: Piras, Giorgio, et al.
Veröffentlicht: (2024)
TabularBench: Benchmarking Adversarial Robustness for Tabular Deep Learning in Real-world Use-cases
von: Simonetto, Thibault, et al.
Veröffentlicht: (2024)
von: Simonetto, Thibault, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025) -
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
von: Nöther, Jonathan, et al.
Veröffentlicht: (2026) -
Policy Teaching via Data Poisoning in Learning from Human Preferences
von: Nika, Andi, et al.
Veröffentlicht: (2025) -
Corruption-Robust Offline Two-Player Zero-Sum Markov Games
von: Nika, Andi, et al.
Veröffentlicht: (2024) -
Corruption Robust Offline Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)