Position: Foundation Models Need Digital Twin Representations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shen, Yiqing, Ding, Hao, Seenivasan, Lalithkumar, Shu, Tianmin, Unberath, Mathias
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909602928394240
author Shen, Yiqing
Ding, Hao
Seenivasan, Lalithkumar
Shu, Tianmin
Unberath, Mathias
author_facet Shen, Yiqing
Ding, Hao
Seenivasan, Lalithkumar
Shu, Tianmin
Unberath, Mathias
contents Current foundation models (FMs) rely on token representations that directly fragment continuous real-world multimodal data into discrete tokens. They limit FMs to learning real-world knowledge and relationships purely through statistical correlation rather than leveraging explicit domain knowledge. Consequently, current FMs struggle with maintaining semantic coherence across modalities, capturing fine-grained spatial-temporal dynamics, and performing causal reasoning. These limitations cannot be overcome by simply scaling up model size or expanding datasets. This position paper argues that the machine learning community should consider digital twin (DT) representations, which are outcome-driven digital representations that serve as building blocks for creating virtual replicas of physical processes, as an alternative to the token representation for building FMs. Finally, we discuss how DT representations can address these challenges by providing physically grounded representations that explicitly encode domain knowledge and preserve the continuous nature of real-world processes.
format Preprint
id arxiv_https___arxiv_org_abs_2505_03798
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Position: Foundation Models Need Digital Twin Representations
Shen, Yiqing
Ding, Hao
Seenivasan, Lalithkumar
Shu, Tianmin
Unberath, Mathias
Machine Learning
Artificial Intelligence
Current foundation models (FMs) rely on token representations that directly fragment continuous real-world multimodal data into discrete tokens. They limit FMs to learning real-world knowledge and relationships purely through statistical correlation rather than leveraging explicit domain knowledge. Consequently, current FMs struggle with maintaining semantic coherence across modalities, capturing fine-grained spatial-temporal dynamics, and performing causal reasoning. These limitations cannot be overcome by simply scaling up model size or expanding datasets. This position paper argues that the machine learning community should consider digital twin (DT) representations, which are outcome-driven digital representations that serve as building blocks for creating virtual replicas of physical processes, as an alternative to the token representation for building FMs. Finally, we discuss how DT representations can address these challenges by providing physically grounded representations that explicitly encode domain knowledge and preserve the continuous nature of real-world processes.
title Position: Foundation Models Need Digital Twin Representations
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.03798