Monotone and Conservative Policy Iteration Beyond the Tabular Case
Fuente:
arXiv
Saved in:
| Main Authors: | Eshwar, S. R., Thoppe, Gugan, Barua, Ananyabrata, Gopalan, Aditya, Dalal, Gal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Does DQN Learn?
by: Gopalan, Aditya, et al.
Published: (2022)
by: Gopalan, Aditya, et al.
Published: (2022)
ElasticFlow: One-Step Physics-Consistent Policy with Elastic Time Horizons for Language-Guided Manipulation
by: Chen, Kewei, et al.
Published: (2026)
by: Chen, Kewei, et al.
Published: (2026)
Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning
by: Chen, Kewei, et al.
Published: (2026)
by: Chen, Kewei, et al.
Published: (2026)
CBX: Python and Julia packages for consensus-based interacting particle methods
by: Bailo, Rafael, et al.
Published: (2024)
by: Bailo, Rafael, et al.
Published: (2024)
Asymptotically Optimal Policies for Weakly Coupled Markov Decision Processes
by: Goldsztajn, Diego, et al.
Published: (2024)
by: Goldsztajn, Diego, et al.
Published: (2024)
PIPHEN: Physical Interaction Prediction with Hamiltonian Energy Networks
by: Chen, Kewei, et al.
Published: (2025)
by: Chen, Kewei, et al.
Published: (2025)
Multi-level meta-reinforcement learning with skill-based curriculum
by: Yang, Sichen, et al.
Published: (2026)
by: Yang, Sichen, et al.
Published: (2026)
LeLaR: The First In-Orbit Demonstration of an AI-Based Satellite Attitude Controller
by: Djebko, Kirill, et al.
Published: (2025)
by: Djebko, Kirill, et al.
Published: (2025)
Operator-Theoretic Foundations and Policy Gradient Methods for General MDPs with Unbounded Costs
by: Gupta, Abhishek, et al.
Published: (2026)
by: Gupta, Abhishek, et al.
Published: (2026)
TOPSIS-like metaheuristic for LABS problem
by: Urbańczyk, Aleksandra, et al.
Published: (2025)
by: Urbańczyk, Aleksandra, et al.
Published: (2025)
Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation
by: Osman, Asim, et al.
Published: (2026)
by: Osman, Asim, et al.
Published: (2026)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
by: Chen, Kewei, et al.
Published: (2025)
by: Chen, Kewei, et al.
Published: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
by: Chen, Kewei, et al.
Published: (2026)
by: Chen, Kewei, et al.
Published: (2026)
Universal Neural Optimal Transport
by: Geuter, Jonathan, et al.
Published: (2022)
by: Geuter, Jonathan, et al.
Published: (2022)
Autonomous AI Agents for Real-Time Affordable Housing Site Selection: Multi-Objective Reinforcement Learning Under Regulatory Constraints
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
Learning to Choose Branching Rules for Nonconvex MINLPs
by: Berthold, Timo, et al.
Published: (2026)
by: Berthold, Timo, et al.
Published: (2026)
Deep Bilinear Koopman Model for Real-Time Vehicle Control in Frenet Frame
by: Abtahi, Mohammad, et al.
Published: (2025)
by: Abtahi, Mohammad, et al.
Published: (2025)
Neural Operators for Mathematical Modeling of Transient Fluid Flow in Subsurface Reservoir Systems
by: Sirota, Daniil D., et al.
Published: (2025)
by: Sirota, Daniil D., et al.
Published: (2025)
Unichain and Aperiodicity are Sufficient for Asymptotic Optimality of Average-Reward Restless Bandits
by: Hong, Yige, et al.
Published: (2024)
by: Hong, Yige, et al.
Published: (2024)
Deep Relaxation of Controlled Stochastic Gradient Descent via Singular Perturbations
by: Bardi, Martino, et al.
Published: (2022)
by: Bardi, Martino, et al.
Published: (2022)
Functional Similarity Metric for Neural Networks: Overcoming Parametric Ambiguity via Activation Region Analysis
by: Hennadii, Kutomanov
Published: (2026)
by: Hennadii, Kutomanov
Published: (2026)
A Note on Stability in Asynchronous Stochastic Approximation without Communication Delays
by: Yu, Huizhen, et al.
Published: (2023)
by: Yu, Huizhen, et al.
Published: (2023)
Piecewise Linear Approximation and PID Control Optimization for Nonlinear Systems
by: Vrabel, Robert
Published: (2025)
by: Vrabel, Robert
Published: (2025)
Bayesian Conservative Policy Optimization (BCPO): A Novel Uncertainty-Calibrated Offline Reinforcement Learning with Credible Lower Bounds
by: Chatterjee, Debashis
Published: (2026)
by: Chatterjee, Debashis
Published: (2026)
Weakly-Coupled Multi-Action Restless Bandits -- Exponential Convergence in Probability
by: Fu, Jing, et al.
Published: (2026)
by: Fu, Jing, et al.
Published: (2026)
Understanding the Impact of Hydro-Reservoirs and Inverters on Frequency-Constrained Operation
by: Aravena, Valeria, et al.
Published: (2025)
by: Aravena, Valeria, et al.
Published: (2025)
Trajectory Optimization and NMPC Tracking for a Fixed Wing UAV in Deep Stall with Perch Landing
by: Nguyen, Huu Thien, et al.
Published: (2022)
by: Nguyen, Huu Thien, et al.
Published: (2022)
Autonomous Cyber Resilience via a Co-Evolutionary Arms Race within a Fortified Digital Twin Sandbox
by: Malikussaid, et al.
Published: (2025)
by: Malikussaid, et al.
Published: (2025)
Sequential, Parallel and Consecutive Hybrid Evolutionary-Swarm Optimization Metaheuristics
by: Urbańczyk, Piotr, et al.
Published: (2025)
by: Urbańczyk, Piotr, et al.
Published: (2025)
Anarchy in the swarm: Testing informed and uninformed diversity-enhancing mechanisms within PSO framework
by: Urbańczyk, Piotr, et al.
Published: (2026)
by: Urbańczyk, Piotr, et al.
Published: (2026)
Decision-Making for Land Conservation: A Derivative-Free Optimization Framework with Nonlinear Inputs
by: Buhler, Cassidy K., et al.
Published: (2023)
by: Buhler, Cassidy K., et al.
Published: (2023)
Enhancing Diversity in Multi-objective Feature Selection
by: Miyandoab, Sevil Zanjani, et al.
Published: (2024)
by: Miyandoab, Sevil Zanjani, et al.
Published: (2024)
State-Dependent Uncertainty Modeling in Robust Optimal Control Problems through Generalized Semi-Infinite Programming
by: Wehbeh, J., et al.
Published: (2025)
by: Wehbeh, J., et al.
Published: (2025)
Socio-cognitive agent-oriented evolutionary algorithm with trust-based optimization
by: Urbańczyk, Aleksandra, et al.
Published: (2025)
by: Urbańczyk, Aleksandra, et al.
Published: (2025)
Limitations of Scalarisation in MORL: A Comparative Study in Discrete Environments
by: Shah, Muhammad Sa'ood, et al.
Published: (2025)
by: Shah, Muhammad Sa'ood, et al.
Published: (2025)
Pre-trained Visual Representations Generalize Where it Matters in Model-Based Reinforcement Learning
by: Jones, Scott, et al.
Published: (2025)
by: Jones, Scott, et al.
Published: (2025)
Model-based Bootstrap of Controlled Markov Chains
by: Su, Ziwei, et al.
Published: (2026)
by: Su, Ziwei, et al.
Published: (2026)
Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting
by: Gess, Benjamin, et al.
Published: (2023)
by: Gess, Benjamin, et al.
Published: (2023)
Variance Reduced Policy Gradient Method for Multi-Objective Reinforcement Learning
by: Guidobene, Davide, et al.
Published: (2025)
by: Guidobene, Davide, et al.
Published: (2025)
Competition-Based Resilience in Distributed Quadratic Optimization
by: Ballotta, Luca, et al.
Published: (2022)
by: Ballotta, Luca, et al.
Published: (2022)
Similar Items
-
Does DQN Learn?
by: Gopalan, Aditya, et al.
Published: (2022) -
ElasticFlow: One-Step Physics-Consistent Policy with Elastic Time Horizons for Language-Guided Manipulation
by: Chen, Kewei, et al.
Published: (2026) -
Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning
by: Chen, Kewei, et al.
Published: (2026) -
CBX: Python and Julia packages for consensus-based interacting particle methods
by: Bailo, Rafael, et al.
Published: (2024) -
Asymptotically Optimal Policies for Weakly Coupled Markov Decision Processes
by: Goldsztajn, Diego, et al.
Published: (2024)