Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
Fuente:
arXiv
Saved in:
| Main Authors: | Eshwar, S. R., Mukherjee, Aniruddha, Saha, Kintan, Agarwal, Krishna, Thoppe, Gugan, Gopalan, Aditya, Dalal, Gal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Monotone and Conservative Policy Iteration Beyond the Tabular Case
by: Eshwar, S. R., et al.
Published: (2025)
by: Eshwar, S. R., et al.
Published: (2025)
Towards Reliable, Uncertainty-Aware Alignment
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
Reinforcement Learning with Quasi-Hyperbolic Discounting
by: Eshwar, S. R., et al.
Published: (2024)
by: Eshwar, S. R., et al.
Published: (2024)
Does DQN Learn?
by: Gopalan, Aditya, et al.
Published: (2022)
by: Gopalan, Aditya, et al.
Published: (2022)
Online Learning of Weakly Coupled MDP Policies for Load Balancing and Auto Scaling
by: Eshwar, S. R., et al.
Published: (2024)
by: Eshwar, S. R., et al.
Published: (2024)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Sequential Attention-based Sampling for Histopathological Analysis
by: G, Tarun, et al.
Published: (2025)
by: G, Tarun, et al.
Published: (2025)
Towards Reliable Alignment: Uncertainty-aware RLHF
by: Banerjee, Debangshu, et al.
Published: (2024)
by: Banerjee, Debangshu, et al.
Published: (2024)
End-to-End and Phase-Level Performance Optimization for Hyperledger Fabric
by: Sollu, Pavan, et al.
Published: (2026)
by: Sollu, Pavan, et al.
Published: (2026)
Teaching Precommitted Agents: Model-Free Policy Evaluation and Control in Quasi-Hyperbolic Discounted MDPs
by: Eshwar, S. R.
Published: (2025)
by: Eshwar, S. R.
Published: (2025)
The Shadow knows: Empirical Distributions of Minimum Spanning Acycles and Persistence Diagrams of Random Complexes
by: Fraiman, Nicolas, et al.
Published: (2020)
by: Fraiman, Nicolas, et al.
Published: (2020)
PlaMo: Plan and Move in Rich 3D Physical Environments
by: Hallak, Assaf, et al.
Published: (2024)
by: Hallak, Assaf, et al.
Published: (2024)
Tree Search-Based Policy Optimization under Stochastic Execution Delay
by: Valensi, David, et al.
Published: (2024)
by: Valensi, David, et al.
Published: (2024)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
by: Banerjee, Debangshu, et al.
Published: (2023)
by: Banerjee, Debangshu, et al.
Published: (2023)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
by: Du, Yihan, et al.
Published: (2024)
by: Du, Yihan, et al.
Published: (2024)
More Test-Time Compute Can Hurt: Overestimation Bias in LLM Beam Search
by: Dalal, Gal, et al.
Published: (2026)
by: Dalal, Gal, et al.
Published: (2026)
Robustness Analysis of POMDP Policies to Observation Perturbations
by: Kraske, Benjamin, et al.
Published: (2026)
by: Kraske, Benjamin, et al.
Published: (2026)
Gradient Boosting Reinforcement Learning
by: Fuhrer, Benjamin, et al.
Published: (2024)
by: Fuhrer, Benjamin, et al.
Published: (2024)
Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
by: Zhao, Zirui, et al.
Published: (2024)
by: Zhao, Zirui, et al.
Published: (2024)
Howard's Policy Iteration is Subexponential for Deterministic Markov Decision Problems with Rewards of Fixed Bit-size and Arbitrary Discount Factor
by: Mukherjee, Dibyangshu, et al.
Published: (2025)
by: Mukherjee, Dibyangshu, et al.
Published: (2025)
Parameter-Free Federated TD Learning with Markov Noise in Heterogeneous Environments
by: Naskar, Ankur, et al.
Published: (2025)
by: Naskar, Ankur, et al.
Published: (2025)
LUMOS: Large User MOdels for User Behavior Prediction
by: Nigam, Dhruv, et al.
Published: (2025)
by: Nigam, Dhruv, et al.
Published: (2025)
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
The Bayesian Geometry of Transformer Attention
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Geometric Scaling of Bayesian Inference in LLMs
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Composing Diffusion Policies for Few-shot Learning of Movement Trajectories
by: Patil, Omkar, et al.
Published: (2024)
by: Patil, Omkar, et al.
Published: (2024)
Greedy Perspectives: Multi-Drone View Planning for Collaborative Perception in Cluttered Environments
by: Suresh, Krishna, et al.
Published: (2023)
by: Suresh, Krishna, et al.
Published: (2023)
Robust Regularized Policy Iteration under Transition Uncertainty
by: Lin, Hongqiang, et al.
Published: (2026)
by: Lin, Hongqiang, et al.
Published: (2026)
Parameter-free Optimal Rates for Nonlinear Semi-Norm Contractions with Applications to $Q$-Learning
by: Naskar, Ankur, et al.
Published: (2025)
by: Naskar, Ankur, et al.
Published: (2025)
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
GOPO: Policy Optimization using Ranked Rewards
by: Choi, Kyuseong, et al.
Published: (2026)
by: Choi, Kyuseong, et al.
Published: (2026)
EDMP: Ensemble-of-costs-guided Diffusion for Motion Planning
by: Saha, Kallol, et al.
Published: (2023)
by: Saha, Kallol, et al.
Published: (2023)
SAM for Robust Mitochondria Instance Segmentation in Fluorescence Microscopy
by: Jadhav, Suyog, et al.
Published: (2026)
by: Jadhav, Suyog, et al.
Published: (2026)
Robustness as Architecture: Designing IQA Models to Withstand Adversarial Perturbations
by: Meleshin, Igor, et al.
Published: (2025)
by: Meleshin, Igor, et al.
Published: (2025)
A Collaborative Multi-Agent Approach to Retrieval-Augmented Generation Across Diverse Data
by: Salve, Aniruddha, et al.
Published: (2024)
by: Salve, Aniruddha, et al.
Published: (2024)
Are Large Language Models Truly Smarter Than Humans?
by: M, Eshwar Reddy, et al.
Published: (2026)
by: M, Eshwar Reddy, et al.
Published: (2026)
Prompting Policies for Multi-step Reasoning and Tool-Use in Black-box LLMs with Iterative Distillation of Experience
by: Sayana, Krishna, et al.
Published: (2026)
by: Sayana, Krishna, et al.
Published: (2026)
Strongly Polynomial Time Complexity of Policy Iteration for $L_\infty$ Robust MDPs
by: Asadi, Ali, et al.
Published: (2026)
by: Asadi, Ali, et al.
Published: (2026)
Action-Inspired Generative Models
by: A., Eshwar R., et al.
Published: (2026)
by: A., Eshwar R., et al.
Published: (2026)
Similar Items
-
Monotone and Conservative Policy Iteration Beyond the Tabular Case
by: Eshwar, S. R., et al.
Published: (2025) -
Towards Reliable, Uncertainty-Aware Alignment
by: Banerjee, Debangshu, et al.
Published: (2025) -
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023) -
Reinforcement Learning with Quasi-Hyperbolic Discounting
by: Eshwar, S. R., et al.
Published: (2024) -
Does DQN Learn?
by: Gopalan, Aditya, et al.
Published: (2022)