Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hsu, YuChe, Wang, AnJui, Ni, TsaiChing, Yang, YuanFu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912819323076608
author Hsu, YuChe
Wang, AnJui
Ni, TsaiChing
Yang, YuanFu
author_facet Hsu, YuChe
Wang, AnJui
Ni, TsaiChing
Yang, YuanFu
contents We propose a Vision-Language Simulation Model (VLSM) that unifies visual and textual understanding to synthesize executable FlexScript from layout sketches and natural-language prompts, enabling cross-modal reasoning for industrial simulation systems. To support this new paradigm, the study constructs the first large-scale dataset for generative digital twins, comprising over 120,000 prompt-sketch-code triplets that enable multimodal learning between textual descriptions, spatial structures, and simulation logic. In parallel, three novel evaluation metrics, Structural Validity Rate (SVR), Parameter Match Rate (PMR), and Execution Success Rate (ESR), are proposed specifically for this task to comprehensively evaluate structural integrity, parameter fidelity, and simulator executability. Through systematic ablation across vision encoders, connectors, and code-pretrained language backbones, the proposed models achieve near-perfect structural accuracy and high execution robustness. This work establishes a foundation for generative digital twins that integrate visual reasoning and language understanding into executable industrial simulation systems. Project page: https://danielhsu2014.github.io/GDT-VLSM-project/
format Preprint
id arxiv_https___arxiv_org_abs_2512_20387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
Hsu, YuChe
Wang, AnJui
Ni, TsaiChing
Yang, YuanFu
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
We propose a Vision-Language Simulation Model (VLSM) that unifies visual and textual understanding to synthesize executable FlexScript from layout sketches and natural-language prompts, enabling cross-modal reasoning for industrial simulation systems. To support this new paradigm, the study constructs the first large-scale dataset for generative digital twins, comprising over 120,000 prompt-sketch-code triplets that enable multimodal learning between textual descriptions, spatial structures, and simulation logic. In parallel, three novel evaluation metrics, Structural Validity Rate (SVR), Parameter Match Rate (PMR), and Execution Success Rate (ESR), are proposed specifically for this task to comprehensively evaluate structural integrity, parameter fidelity, and simulator executability. Through systematic ablation across vision encoders, connectors, and code-pretrained language backbones, the proposed models achieve near-perfect structural accuracy and high execution robustness. This work establishes a foundation for generative digital twins that integrate visual reasoning and language understanding into executable industrial simulation systems. Project page: https://danielhsu2014.github.io/GDT-VLSM-project/
title Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.20387