Geometry of Drifting MDPs with Path-Integral Stability Certificates
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zuyuan, Imani, Mahdi, Lan, Tian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Metric-Gradient Projection for Stable Multi-Agent Policy Learning
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
Manifold-Constrained Energy-Based Transition Models for Offline Reinforcement Learning
by: Fang, Zeyu, et al.
Published: (2026)
by: Fang, Zeyu, et al.
Published: (2026)
Interactive Critique-Revision Training for Reliable Structured LLM Generation
by: Yu, Fei Xu, et al.
Published: (2026)
by: Yu, Fei Xu, et al.
Published: (2026)
Matrix-Space Reinforcement Learning for Reusing Local Transition Geometry
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees
by: Zhang, Zuyuan, et al.
Published: (2025)
by: Zhang, Zuyuan, et al.
Published: (2025)
Operator-Guided Invariance Learning for Continuous Reinforcement Learning
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search
by: Tang, Sizhe, et al.
Published: (2026)
by: Tang, Sizhe, et al.
Published: (2026)
Global Optimization on Graph-Structured Data via Gaussian Processes with Spectral Representations
by: Hong, Shu, et al.
Published: (2025)
by: Hong, Shu, et al.
Published: (2025)
Collaborative AI Teaming in Unknown Environments via Active Goal Deduction
by: Zhang, Zuyuan, et al.
Published: (2024)
by: Zhang, Zuyuan, et al.
Published: (2024)
Lipschitz Lifelong Monte Carlo Tree Search for Mastering Non-Stationary Tasks
by: Zhang, Zuyuan, et al.
Published: (2025)
by: Zhang, Zuyuan, et al.
Published: (2025)
Cooperative Backdoor Attack in Decentralized Reinforcement Learning with Theoretical Guarantee
by: Gao, Mengtong, et al.
Published: (2024)
by: Gao, Mengtong, et al.
Published: (2024)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Learning Policy Committees for Effective Personalization in MDPs with Diverse Tasks
by: Ge, Luise, et al.
Published: (2025)
by: Ge, Luise, et al.
Published: (2025)
Low-Rank MDPs with Continuous Action Spaces
by: Bennett, Andrew, et al.
Published: (2023)
by: Bennett, Andrew, et al.
Published: (2023)
Efficient Solution and Learning of Robust Factored MDPs
by: Schnitzer, Yannik, et al.
Published: (2025)
by: Schnitzer, Yannik, et al.
Published: (2025)
Navigating Explanatory Multiverse Through Counterfactual Path Geometry
by: Sokol, Kacper, et al.
Published: (2023)
by: Sokol, Kacper, et al.
Published: (2023)
SR-Reward: Taking The Path More Traveled
by: Azad, Seyed Mahdi B., et al.
Published: (2025)
by: Azad, Seyed Mahdi B., et al.
Published: (2025)
Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs
by: Wendland, Joshua, et al.
Published: (2026)
by: Wendland, Joshua, et al.
Published: (2026)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
by: Shakerinava, Mehran, et al.
Published: (2025)
by: Shakerinava, Mehran, et al.
Published: (2025)
Second-Order Convergence in Private Stochastic Non-Convex Optimization
by: Tao, Youming, et al.
Published: (2025)
by: Tao, Youming, et al.
Published: (2025)
$n$-Musketeers: Reinforcement Learning Shapes Collaboration Among Language Models
by: Masukawa, Ryozo, et al.
Published: (2026)
by: Masukawa, Ryozo, et al.
Published: (2026)
Exact Attention Sensitivity and the Geometry of Transformer Stability
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
by: Shah, Anvay, et al.
Published: (2026)
by: Shah, Anvay, et al.
Published: (2026)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023)
by: Agnihotri, Akhil, et al.
Published: (2023)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
by: Zimmer, Matthieu, et al.
Published: (2025)
by: Zimmer, Matthieu, et al.
Published: (2025)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
Solving Multi-Model MDPs by Coordinate Ascent and Dynamic Programming
by: Su, Xihong, et al.
Published: (2024)
by: Su, Xihong, et al.
Published: (2024)
Risk-averse Total-reward MDPs with ERM and EVaR
by: Su, Xihong, et al.
Published: (2024)
by: Su, Xihong, et al.
Published: (2024)
On the Convergence of Monte Carlo UCB for Random-Length Episodic MDPs
by: Dong, Zixuan, et al.
Published: (2022)
by: Dong, Zixuan, et al.
Published: (2022)
Drift-Based Dataset Stability Benchmark
by: Soukup, Dominik, et al.
Published: (2025)
by: Soukup, Dominik, et al.
Published: (2025)
Ensuring Safety in an Uncertain Environment: Constrained MDPs via Stochastic Thresholds
by: Zuo, Qian, et al.
Published: (2025)
by: Zuo, Qian, et al.
Published: (2025)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
by: Banerjee, Debangshu, et al.
Published: (2023)
by: Banerjee, Debangshu, et al.
Published: (2023)
A Testable Certificate for Constant Collapse in Teacher-Guided VAEs
by: Zhang, Zegu, et al.
Published: (2026)
by: Zhang, Zegu, et al.
Published: (2026)
A Simplex Witness Certificate for Constant Collapse in Variational Autoencoders
by: Zhang, Zegu, et al.
Published: (2026)
by: Zhang, Zegu, et al.
Published: (2026)
Physics-Informed Causal MDPs for Sequential Constraint Repair in Engineering Simulation Pipelines
by: Qiao, Chuhan
Published: (2026)
by: Qiao, Chuhan
Published: (2026)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2025)
by: Hong, Kihyuk, et al.
Published: (2025)
DeepAveragers: Offline Reinforcement Learning by Solving Derived Non-Parametric MDPs
by: Shrestha, Aayam, et al.
Published: (2020)
by: Shrestha, Aayam, et al.
Published: (2020)
Projection by Convolution: Optimal Sample Complexity for Reinforcement Learning in Continuous-Space MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Similar Items
-
Metric-Gradient Projection for Stable Multi-Agent Policy Learning
by: Zhang, Zuyuan, et al.
Published: (2026) -
Manifold-Constrained Energy-Based Transition Models for Offline Reinforcement Learning
by: Fang, Zeyu, et al.
Published: (2026) -
Interactive Critique-Revision Training for Reliable Structured LLM Generation
by: Yu, Fei Xu, et al.
Published: (2026) -
Matrix-Space Reinforcement Learning for Reusing Local Transition Geometry
by: Zhang, Zuyuan, et al.
Published: (2026) -
Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics
by: Zhang, Zuyuan, et al.
Published: (2026)