Toward a Metrology for Artificial Intelligence: Hidden-Rule Environments and Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mathew, Christo, Wang, Wentian, Feldman, Jacob, Gallos, Lazaros K., Kantor, Paul B., Menkov, Vladimir, Wang, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912665535774720
author Mathew, Christo
Wang, Wentian
Feldman, Jacob
Gallos, Lazaros K.
Kantor, Paul B.
Menkov, Vladimir
Wang, Hao
author_facet Mathew, Christo
Wang, Wentian
Feldman, Jacob
Gallos, Lazaros K.
Kantor, Paul B.
Menkov, Vladimir
Wang, Hao
contents We investigate reinforcement learning in the Game Of Hidden Rules (GOHR) environment, a complex puzzle in which an agent must infer and execute hidden rules to clear a 6$\times$6 board by placing game pieces into buckets. We explore two state representation strategies, namely Feature-Centric (FC) and Object-Centric (OC), and employ a Transformer-based Advantage Actor-Critic (A2C) algorithm for training. The agent has access only to partial observations and must simultaneously infer the governing rule and learn the optimal policy through experience. We evaluate our models across multiple rule-based and trial-list-based experimental setups, analyzing transfer effects and the impact of representation on learning efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06213
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toward a Metrology for Artificial Intelligence: Hidden-Rule Environments and Reinforcement Learning
Mathew, Christo
Wang, Wentian
Feldman, Jacob
Gallos, Lazaros K.
Kantor, Paul B.
Menkov, Vladimir
Wang, Hao
Machine Learning
Artificial Intelligence
We investigate reinforcement learning in the Game Of Hidden Rules (GOHR) environment, a complex puzzle in which an agent must infer and execute hidden rules to clear a 6$\times$6 board by placing game pieces into buckets. We explore two state representation strategies, namely Feature-Centric (FC) and Object-Centric (OC), and employ a Transformer-based Advantage Actor-Critic (A2C) algorithm for training. The agent has access only to partial observations and must simultaneously infer the governing rule and learn the optimal policy through experience. We evaluate our models across multiple rule-based and trial-list-based experimental setups, analyzing transfer effects and the impact of representation on learning efficiency.
title Toward a Metrology for Artificial Intelligence: Hidden-Rule Environments and Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.06213