Deep Reinforcement Learning Agents are not even close to Human Intelligence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Delfosse, Quentin, Blüml, Jannis, Tatai, Fabian, Vincent, Théo, Gregori, Bjarne, Dillies, Elisabeth, Peters, Jan, Rothkopf, Constantin, Kersting, Kristian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913862322749440
author Delfosse, Quentin
Blüml, Jannis
Tatai, Fabian
Vincent, Théo
Gregori, Bjarne
Dillies, Elisabeth
Peters, Jan
Rothkopf, Constantin
Kersting, Kristian
author_facet Delfosse, Quentin
Blüml, Jannis
Tatai, Fabian
Vincent, Théo
Gregori, Bjarne
Dillies, Elisabeth
Peters, Jan
Rothkopf, Constantin
Kersting, Kristian
contents Deep reinforcement learning (RL) agents achieve impressive results in a wide variety of tasks, but they lack zero-shot adaptation capabilities. While most robustness evaluations focus on tasks complexifications, for which human also struggle to maintain performances, no evaluation has been performed on tasks simplifications. To tackle this issue, we introduce HackAtari, a set of task variations of the Arcade Learning Environments. We use it to demonstrate that, contrary to humans, RL agents systematically exhibit huge performance drops on simpler versions of their training tasks, uncovering agents' consistent reliance on shortcuts. Our analysis across multiple algorithms and architectures highlights the persistent gap between RL agents and human behavioral intelligence, underscoring the need for new benchmarks and methodologies that enforce systematic generalization testing beyond static evaluation protocols. Training and testing in the same environment is not enough to obtain agents equipped with human-like intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21731
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deep Reinforcement Learning Agents are not even close to Human Intelligence
Delfosse, Quentin
Blüml, Jannis
Tatai, Fabian
Vincent, Théo
Gregori, Bjarne
Dillies, Elisabeth
Peters, Jan
Rothkopf, Constantin
Kersting, Kristian
Machine Learning
Artificial Intelligence
Deep reinforcement learning (RL) agents achieve impressive results in a wide variety of tasks, but they lack zero-shot adaptation capabilities. While most robustness evaluations focus on tasks complexifications, for which human also struggle to maintain performances, no evaluation has been performed on tasks simplifications. To tackle this issue, we introduce HackAtari, a set of task variations of the Arcade Learning Environments. We use it to demonstrate that, contrary to humans, RL agents systematically exhibit huge performance drops on simpler versions of their training tasks, uncovering agents' consistent reliance on shortcuts. Our analysis across multiple algorithms and architectures highlights the persistent gap between RL agents and human behavioral intelligence, underscoring the need for new benchmarks and methodologies that enforce systematic generalization testing beyond static evaluation protocols. Training and testing in the same environment is not enough to obtain agents equipped with human-like intelligence.
title Deep Reinforcement Learning Agents are not even close to Human Intelligence
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.21731