Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918378112811008 |
|---|---|
| author | Legel, Lance Huang, Qin Voelker, Brandon Neamati, Daniel Johnson, Patrick Alan Bastani, Favyen Rose, Jeff Hennessy, James Ryan Guralnick, Robert Soltis, Douglas Soltis, Pamela Wang, Shaowen |
| author_facet | Legel, Lance Huang, Qin Voelker, Brandon Neamati, Daniel Johnson, Patrick Alan Bastani, Favyen Rose, Jeff Hennessy, James Ryan Guralnick, Robert Soltis, Douglas Soltis, Pamela Wang, Shaowen |
| contents | We present DeepEarth, a self-supervised multi-modal world model with Earth4D, a novel planetary-scale 4D space-time positional encoder. Earth4D extends 3D multi-resolution hash encoding to include time, efficiently scaling across the planet over centuries with sub-meter, sub-second precision. Multi-modal encoders (e.g. vision-language models) are fused with Earth4D embeddings and trained via masked reconstruction. We demonstrate Earth4D's expressive power by achieving state-of-the-art performance on an ecological forecasting benchmark. Earth4D with learnable hash probing surpasses a multi-modal foundation model pre-trained on substantially more data. Access open source code and download models at: https://github.com/legel/deepearth |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_07039 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding Legel, Lance Huang, Qin Voelker, Brandon Neamati, Daniel Johnson, Patrick Alan Bastani, Favyen Rose, Jeff Hennessy, James Ryan Guralnick, Robert Soltis, Douglas Soltis, Pamela Wang, Shaowen Artificial Intelligence I.2; I.6; J.2; J.3 We present DeepEarth, a self-supervised multi-modal world model with Earth4D, a novel planetary-scale 4D space-time positional encoder. Earth4D extends 3D multi-resolution hash encoding to include time, efficiently scaling across the planet over centuries with sub-meter, sub-second precision. Multi-modal encoders (e.g. vision-language models) are fused with Earth4D embeddings and trained via masked reconstruction. We demonstrate Earth4D's expressive power by achieving state-of-the-art performance on an ecological forecasting benchmark. Earth4D with learnable hash probing surpasses a multi-modal foundation model pre-trained on substantially more data. Access open source code and download models at: https://github.com/legel/deepearth |
| title | Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding |
| topic | Artificial Intelligence I.2; I.6; J.2; J.3 |
| url | https://arxiv.org/abs/2603.07039 |