The Pursuit of Diversity: Multi-Objective Testing of Deep Reinforcement Learning Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bartlett, Antony, Liem, Cynthia, Panichella, Annibale
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911215152791552
author Bartlett, Antony
Liem, Cynthia
Panichella, Annibale
author_facet Bartlett, Antony
Liem, Cynthia
Panichella, Annibale
contents Testing deep reinforcement learning (DRL) agents in safety-critical domains requires discovering diverse failure scenarios. Existing tools such as INDAGO rely on single-objective optimization focused solely on maximizing failure counts, but this does not ensure discovered scenarios are diverse or reveal distinct error types. We introduce INDAGO-Nexus, a multi-objective search approach that jointly optimizes for failure likelihood and test scenario diversity using multi-objective evolutionary algorithms with multiple diversity metrics and Pareto front selection strategies. We evaluated INDAGO-Nexus on three DRL agents: humanoid walker, self-driving car, and parking agent. On average, INDAGO-Nexus discovers up to 83% and 40% more unique failures (test effectiveness) than INDAGO in the SDC and Parking scenarios, respectively, while reducing time-to-failure by up to 67% across all agents.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14727
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Pursuit of Diversity: Multi-Objective Testing of Deep Reinforcement Learning Agents
Bartlett, Antony
Liem, Cynthia
Panichella, Annibale
Machine Learning
Testing deep reinforcement learning (DRL) agents in safety-critical domains requires discovering diverse failure scenarios. Existing tools such as INDAGO rely on single-objective optimization focused solely on maximizing failure counts, but this does not ensure discovered scenarios are diverse or reveal distinct error types. We introduce INDAGO-Nexus, a multi-objective search approach that jointly optimizes for failure likelihood and test scenario diversity using multi-objective evolutionary algorithms with multiple diversity metrics and Pareto front selection strategies. We evaluated INDAGO-Nexus on three DRL agents: humanoid walker, self-driving car, and parking agent. On average, INDAGO-Nexus discovers up to 83% and 40% more unique failures (test effectiveness) than INDAGO in the SDC and Parking scenarios, respectively, while reducing time-to-failure by up to 67% across all agents.
title The Pursuit of Diversity: Multi-Objective Testing of Deep Reinforcement Learning Agents
topic Machine Learning
url https://arxiv.org/abs/2510.14727