The Observer Effect in World Models: Invasive Adaptation Corrupts Latent Physics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Internò, Christian, Yamaguchi, Jumpei, Amdahl-Culleton, Loren, Olhofer, Markus, Klindt, David, Hammer, Barbara
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908830193942528
author Internò, Christian
Yamaguchi, Jumpei
Amdahl-Culleton, Loren
Olhofer, Markus
Klindt, David
Hammer, Barbara
author_facet Internò, Christian
Yamaguchi, Jumpei
Amdahl-Culleton, Loren
Olhofer, Markus
Klindt, David
Hammer, Barbara
contents Determining whether neural models internalize physical laws as world models, rather than exploiting statistical shortcuts, remains challenging, especially under out-of-distribution (OOD) shifts. Standard evaluations often test latent capability via downstream adaptation (e.g., fine-tuning or high-capacity probes), but such interventions can change the representations being measured and thus confound what was learned during self-supervised learning (SSL). We propose a non-invasive evaluation protocol, PhyIP. We test whether physical quantities are linearly decodable from frozen representations, motivated by the linear representation hypothesis. Across fluid dynamics and orbital mechanics, we find that when SSL achieves low error, latent structure becomes linearly accessible. PhyIP recovers internal energy and Newtonian inverse-square scaling on OOD tests (e.g., $ρ> 0.90$). In contrast, adaptation-based evaluations can collapse this structure ($ρ\approx 0.05$). These findings suggest that adaptation-based evaluation can obscure latent structures and that low-capacity probes offer a more accurate evaluation of physical world models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_12218
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Observer Effect in World Models: Invasive Adaptation Corrupts Latent Physics
Internò, Christian
Yamaguchi, Jumpei
Amdahl-Culleton, Loren
Olhofer, Markus
Klindt, David
Hammer, Barbara
Machine Learning
Artificial Intelligence
Determining whether neural models internalize physical laws as world models, rather than exploiting statistical shortcuts, remains challenging, especially under out-of-distribution (OOD) shifts. Standard evaluations often test latent capability via downstream adaptation (e.g., fine-tuning or high-capacity probes), but such interventions can change the representations being measured and thus confound what was learned during self-supervised learning (SSL). We propose a non-invasive evaluation protocol, PhyIP. We test whether physical quantities are linearly decodable from frozen representations, motivated by the linear representation hypothesis. Across fluid dynamics and orbital mechanics, we find that when SSL achieves low error, latent structure becomes linearly accessible. PhyIP recovers internal energy and Newtonian inverse-square scaling on OOD tests (e.g., $ρ> 0.90$). In contrast, adaptation-based evaluations can collapse this structure ($ρ\approx 0.05$). These findings suggest that adaptation-based evaluation can obscure latent structures and that low-capacity probes offer a more accurate evaluation of physical world models.
title The Observer Effect in World Models: Invasive Adaptation Corrupts Latent Physics
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.12218