Neural Value Iteration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: You, Yang, Çakır, Ufuk, Schutz, Alex, Hawes, Nick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915863694671872
author You, Yang
Çakır, Ufuk
Schutz, Alex
Hawes, Nick
author_facet You, Yang
Çakır, Ufuk
Schutz, Alex
Hawes, Nick
contents The value function of a POMDP exhibits the piecewise-linear-convex (PWLC) property and can be represented as a finite set of hyperplanes, known as $α$-vectors. Most state-of-the-art POMDP solvers (offline planners) follow the point-based value iteration scheme, which performs Bellman backups on $α$-vectors at reachable belief points until convergence. However, since each $α$-vector is $|S|$-dimensional, these methods quickly become intractable for large-scale problems due to the prohibitive computational cost of Bellman backups. In this work, we demonstrate that the PWLC property allows a POMDP's value function to be alternatively represented as a finite set of neural networks. This insight enables a novel POMDP planning algorithm called \emph{Neural Value Iteration}, which combines the generalization capability of neural networks with the classical value iteration framework. Our approach achieves near-optimal solutions even in extremely large POMDPs that are intractable for existing offline solvers.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08825
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Neural Value Iteration
You, Yang
Çakır, Ufuk
Schutz, Alex
Hawes, Nick
Artificial Intelligence
The value function of a POMDP exhibits the piecewise-linear-convex (PWLC) property and can be represented as a finite set of hyperplanes, known as $α$-vectors. Most state-of-the-art POMDP solvers (offline planners) follow the point-based value iteration scheme, which performs Bellman backups on $α$-vectors at reachable belief points until convergence. However, since each $α$-vector is $|S|$-dimensional, these methods quickly become intractable for large-scale problems due to the prohibitive computational cost of Bellman backups. In this work, we demonstrate that the PWLC property allows a POMDP's value function to be alternatively represented as a finite set of neural networks. This insight enables a novel POMDP planning algorithm called \emph{Neural Value Iteration}, which combines the generalization capability of neural networks with the classical value iteration framework. Our approach achieves near-optimal solutions even in extremely large POMDPs that are intractable for existing offline solvers.
title Neural Value Iteration
topic Artificial Intelligence
url https://arxiv.org/abs/2511.08825