Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Seo, SeungWon, Han, DongHeun, Noh, SeongRae, Kang, HyeongYeop
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913134517682176
author Seo, SeungWon
Han, DongHeun
Noh, SeongRae
Kang, HyeongYeop
author_facet Seo, SeungWon
Han, DongHeun
Noh, SeongRae
Kang, HyeongYeop
contents Executable world models can be read, edited, executed, and reused for planning, but only if the program captures the environment's transition law rather than semantic shortcuts in its surface vocabulary. We study online executable world-model learning under prior misalignment, where an agent must induce state-dependent dynamics from interaction evidence alone, without rule descriptions, reward signals, or trustworthy lexical priors. We introduce Alice, a closed-loop system that treats failed candidate updates as structural signal: when a candidate explains a new transition but loses previously explained ones, the preservation conflict reveals dynamics that the current program had conflated. Alice refines these conflicts into hypothesis classes that both provide compact, class-stratified preservation counterexamples for update and guide frontier exploration toward transitions that are novel and underrepresented with respect to the current program. We evaluate Alice on Baba in Wonderland, a prior-misaligned variant of Baba Is You that preserves simulator dynamics while replacing semantically meaningful rule-property labels with unrelated words. Experiments show that Alice substantially improves executable world-model learning under prior misalignment, and ablations show that both class refinement and class-aware exploration contribute.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16725
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models
Seo, SeungWon
Han, DongHeun
Noh, SeongRae
Kang, HyeongYeop
Artificial Intelligence
Executable world models can be read, edited, executed, and reused for planning, but only if the program captures the environment's transition law rather than semantic shortcuts in its surface vocabulary. We study online executable world-model learning under prior misalignment, where an agent must induce state-dependent dynamics from interaction evidence alone, without rule descriptions, reward signals, or trustworthy lexical priors. We introduce Alice, a closed-loop system that treats failed candidate updates as structural signal: when a candidate explains a new transition but loses previously explained ones, the preservation conflict reveals dynamics that the current program had conflated. Alice refines these conflicts into hypothesis classes that both provide compact, class-stratified preservation counterexamples for update and guide frontier exploration toward transitions that are novel and underrepresented with respect to the current program. We evaluate Alice on Baba in Wonderland, a prior-misaligned variant of Baba Is You that preserves simulator dynamics while replacing semantically meaningful rule-property labels with unrelated words. Experiments show that Alice substantially improves executable world-model learning under prior misalignment, and ablations show that both class refinement and class-aware exploration contribute.
title Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models
topic Artificial Intelligence
url https://arxiv.org/abs/2605.16725