Saved in:
Bibliographic Details
Main Authors: Gavenski, Nathan, Rodrigues, Odinaldo
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.24784
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918150599081984
author Gavenski, Nathan
Rodrigues, Odinaldo
author_facet Gavenski, Nathan
Rodrigues, Odinaldo
contents Imitation learning benchmarks often lack sufficient variation between training and evaluation, limiting meaningful generalisation assessment. We introduce Labyrinth, a benchmarking environment designed to test generalisation with precise control over structure, start and goal positions, and task complexity. It enables verifiably distinct training, evaluation, and test settings. Labyrinth provides a discrete, fully observable state space and known optimal actions, supporting interpretability and fine-grained evaluation. Its flexible setup allows targeted testing of generalisation factors and includes variants like partial observability, key-and-door tasks, and ice-floor hazards. By enabling controlled, reproducible experiments, Labyrinth advances the evaluation of generalisation in imitation learning and provides a valuable tool for developing more robust agents.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24784
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quantifying Generalisation in Imitation Learning
Gavenski, Nathan
Rodrigues, Odinaldo
Machine Learning
Artificial Intelligence
Imitation learning benchmarks often lack sufficient variation between training and evaluation, limiting meaningful generalisation assessment. We introduce Labyrinth, a benchmarking environment designed to test generalisation with precise control over structure, start and goal positions, and task complexity. It enables verifiably distinct training, evaluation, and test settings. Labyrinth provides a discrete, fully observable state space and known optimal actions, supporting interpretability and fine-grained evaluation. Its flexible setup allows targeted testing of generalisation factors and includes variants like partial observability, key-and-door tasks, and ice-floor hazards. By enabling controlled, reproducible experiments, Labyrinth advances the evaluation of generalisation in imitation learning and provides a valuable tool for developing more robust agents.
title Quantifying Generalisation in Imitation Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.24784