LyTimeT: Towards Robust and Interpretable State-Variable Discovery

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yu, Kuai, Su, Crystal, Liu, Xiang, Goldfeder, Judah, Shao, Mingyuan, Lipson, Hod
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917034168680448
author Yu, Kuai
Su, Crystal
Liu, Xiang
Goldfeder, Judah
Shao, Mingyuan
Lipson, Hod
author_facet Yu, Kuai
Su, Crystal
Liu, Xiang
Goldfeder, Judah
Shao, Mingyuan
Lipson, Hod
contents Extracting the true dynamical variables of a system from high-dimensional video is challenging due to distracting visual factors such as background motion, occlusions, and texture changes. We propose LyTimeT, a two-phase framework for interpretable variable extraction that learns robust and stable latent representations of dynamical systems. In Phase 1, LyTimeT employs a spatio-temporal TimeSformer-based autoencoder that uses global attention to focus on dynamically relevant regions while suppressing nuisance variation, enabling distraction-robust latent state learning and accurate long-horizon video prediction. In Phase 2, we probe the learned latent space, select the most physically meaningful dimensions using linear correlation analysis, and refine the transition dynamics with a Lyapunov-based stability regularizer to enforce contraction and reduce error accumulation during roll-outs. Experiments on five synthetic benchmarks and four real-world dynamical systems, including chaotic phenomena, show that LyTimeT achieves mutual information and intrinsic dimension estimates closest to ground truth, remains invariant under background perturbations, and delivers the lowest analytical mean squared error among CNN-based (TIDE) and transformer-only baselines. Our results demonstrate that combining spatio-temporal attention with stability constraints yields predictive models that are not only accurate but also physically interpretable.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19716
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LyTimeT: Towards Robust and Interpretable State-Variable Discovery
Yu, Kuai
Su, Crystal
Liu, Xiang
Goldfeder, Judah
Shao, Mingyuan
Lipson, Hod
Computer Vision and Pattern Recognition
Extracting the true dynamical variables of a system from high-dimensional video is challenging due to distracting visual factors such as background motion, occlusions, and texture changes. We propose LyTimeT, a two-phase framework for interpretable variable extraction that learns robust and stable latent representations of dynamical systems. In Phase 1, LyTimeT employs a spatio-temporal TimeSformer-based autoencoder that uses global attention to focus on dynamically relevant regions while suppressing nuisance variation, enabling distraction-robust latent state learning and accurate long-horizon video prediction. In Phase 2, we probe the learned latent space, select the most physically meaningful dimensions using linear correlation analysis, and refine the transition dynamics with a Lyapunov-based stability regularizer to enforce contraction and reduce error accumulation during roll-outs. Experiments on five synthetic benchmarks and four real-world dynamical systems, including chaotic phenomena, show that LyTimeT achieves mutual information and intrinsic dimension estimates closest to ground truth, remains invariant under background perturbations, and delivers the lowest analytical mean squared error among CNN-based (TIDE) and transformer-only baselines. Our results demonstrate that combining spatio-temporal attention with stability constraints yields predictive models that are not only accurate but also physically interpretable.
title LyTimeT: Towards Robust and Interpretable State-Variable Discovery
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.19716