On the Ability of Transformers to Verify Plans

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sarrof, Yash, Du, Yupei, Stein, Katharina, Koller, Alexander, Thiébaux, Sylvie, Hahn, Michael
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910061333315584
author Sarrof, Yash
Du, Yupei
Stein, Katharina
Koller, Alexander
Thiébaux, Sylvie
Hahn, Michael
author_facet Sarrof, Yash
Du, Yupei
Stein, Katharina
Koller, Alexander
Thiébaux, Sylvie
Hahn, Michael
contents Transformers have shown inconsistent success in AI planning tasks, and theoretical understanding of when generalization should be expected has been limited. We take important steps towards addressing this gap by analyzing the ability of decoder-only models to verify whether a given plan correctly solves a given planning instance. To analyse the general setting where the number of objects -- and thus the effective input alphabet -- grows at test time, we introduce C*-RASP, an extension of C-RASP designed to establish length generalization guarantees for transformers under the simultaneous growth in sequence length and vocabulary size. Our results identify a large class of classical planning domains for which transformers can provably learn to verify long plans, and structural properties that significantly affects the learnability of length generalizable solutions. Empirical experiments corroborate our theory.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19954
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On the Ability of Transformers to Verify Plans
Sarrof, Yash
Du, Yupei
Stein, Katharina
Koller, Alexander
Thiébaux, Sylvie
Hahn, Michael
Artificial Intelligence
Computation and Language
Machine Learning
Transformers have shown inconsistent success in AI planning tasks, and theoretical understanding of when generalization should be expected has been limited. We take important steps towards addressing this gap by analyzing the ability of decoder-only models to verify whether a given plan correctly solves a given planning instance. To analyse the general setting where the number of objects -- and thus the effective input alphabet -- grows at test time, we introduce C*-RASP, an extension of C-RASP designed to establish length generalization guarantees for transformers under the simultaneous growth in sequence length and vocabulary size. Our results identify a large class of classical planning domains for which transformers can provably learn to verify long plans, and structural properties that significantly affects the learnability of length generalizable solutions. Empirical experiments corroborate our theory.
title On the Ability of Transformers to Verify Plans
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2603.19954