Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Legel, Lance, Huang, Qin, Voelker, Brandon, Neamati, Daniel, Johnson, Patrick Alan, Bastani, Favyen, Rose, Jeff, Hennessy, James Ryan, Guralnick, Robert, Soltis, Douglas, Soltis, Pamela, Wang, Shaowen
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918378112811008
author Legel, Lance
Huang, Qin
Voelker, Brandon
Neamati, Daniel
Johnson, Patrick Alan
Bastani, Favyen
Rose, Jeff
Hennessy, James Ryan
Guralnick, Robert
Soltis, Douglas
Soltis, Pamela
Wang, Shaowen
author_facet Legel, Lance
Huang, Qin
Voelker, Brandon
Neamati, Daniel
Johnson, Patrick Alan
Bastani, Favyen
Rose, Jeff
Hennessy, James Ryan
Guralnick, Robert
Soltis, Douglas
Soltis, Pamela
Wang, Shaowen
contents We present DeepEarth, a self-supervised multi-modal world model with Earth4D, a novel planetary-scale 4D space-time positional encoder. Earth4D extends 3D multi-resolution hash encoding to include time, efficiently scaling across the planet over centuries with sub-meter, sub-second precision. Multi-modal encoders (e.g. vision-language models) are fused with Earth4D embeddings and trained via masked reconstruction. We demonstrate Earth4D's expressive power by achieving state-of-the-art performance on an ecological forecasting benchmark. Earth4D with learnable hash probing surpasses a multi-modal foundation model pre-trained on substantially more data. Access open source code and download models at: https://github.com/legel/deepearth
format Preprint
id arxiv_https___arxiv_org_abs_2603_07039
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding
Legel, Lance
Huang, Qin
Voelker, Brandon
Neamati, Daniel
Johnson, Patrick Alan
Bastani, Favyen
Rose, Jeff
Hennessy, James Ryan
Guralnick, Robert
Soltis, Douglas
Soltis, Pamela
Wang, Shaowen
Artificial Intelligence
I.2; I.6; J.2; J.3
We present DeepEarth, a self-supervised multi-modal world model with Earth4D, a novel planetary-scale 4D space-time positional encoder. Earth4D extends 3D multi-resolution hash encoding to include time, efficiently scaling across the planet over centuries with sub-meter, sub-second precision. Multi-modal encoders (e.g. vision-language models) are fused with Earth4D embeddings and trained via masked reconstruction. We demonstrate Earth4D's expressive power by achieving state-of-the-art performance on an ecological forecasting benchmark. Earth4D with learnable hash probing surpasses a multi-modal foundation model pre-trained on substantially more data. Access open source code and download models at: https://github.com/legel/deepearth
title Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding
topic Artificial Intelligence
I.2; I.6; J.2; J.3
url https://arxiv.org/abs/2603.07039