Transformer-like Inference from Optimal Control

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kudre, Aditya, Chang, Heng-Sheng, Mehta, Prashant G.
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910223510274048
author Kudre, Aditya
Chang, Heng-Sheng
Mehta, Prashant G.
author_facet Kudre, Aditya
Chang, Heng-Sheng
Mehta, Prashant G.
contents Decoder-only transformers compute the conditional probability of the next token from a sequence of past observations. This paper derives, from first principles, inference architectures that solve the same prediction problem - and in doing so, recovers transformer-like layer operations as a consequence of optimal control theory. The framework is developed for two model classes: a nonlinear model of discrete-valued processes, directly motivated by the transformer, and a linear Gaussian model as a tractable baseline. For both model classes, the prediction objective is reformulated as an optimal control problem whose solution yields an explicit inference algorithm, the dual filter, with a layer structure that mirrors the layer structure of a decoder-only transformer. Numerical experiments provide a comparison of the optimal control to attention weights from a trained transformer. These experiments reveal that when the embedding dimension is insufficient, the transformer implicitly exploits non-Markovian structure.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15608
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Transformer-like Inference from Optimal Control
Kudre, Aditya
Chang, Heng-Sheng
Mehta, Prashant G.
Machine Learning
Systems and Control
Decoder-only transformers compute the conditional probability of the next token from a sequence of past observations. This paper derives, from first principles, inference architectures that solve the same prediction problem - and in doing so, recovers transformer-like layer operations as a consequence of optimal control theory. The framework is developed for two model classes: a nonlinear model of discrete-valued processes, directly motivated by the transformer, and a linear Gaussian model as a tractable baseline. For both model classes, the prediction objective is reformulated as an optimal control problem whose solution yields an explicit inference algorithm, the dual filter, with a layer structure that mirrors the layer structure of a decoder-only transformer. Numerical experiments provide a comparison of the optimal control to attention weights from a trained transformer. These experiments reveal that when the embedding dimension is insufficient, the transformer implicitly exploits non-Markovian structure.
title Transformer-like Inference from Optimal Control
topic Machine Learning
Systems and Control
url https://arxiv.org/abs/2605.15608