A View of the Certainty-Equivalence Method for PAC RL as an Application of the Trajectory Tree Method
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kalyanakrishnan, Shivaram, Shah, Sheel, Guguloth, Santhosh Kumar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
von: Shah, Anvay, et al.
Veröffentlicht: (2026)
von: Shah, Anvay, et al.
Veröffentlicht: (2026)
Using Common Random Numbers for Simulation-based Planning with Rollouts
von: Yadav, Sandarbh, et al.
Veröffentlicht: (2026)
von: Yadav, Sandarbh, et al.
Veröffentlicht: (2026)
Concentration Bounds for Optimized Certainty Equivalent Risk Estimation
von: Ghosh, Ayon, et al.
Veröffentlicht: (2024)
von: Ghosh, Ayon, et al.
Veröffentlicht: (2024)
Risk-Averse Constrained Reinforcement Learning with Optimized Certainty Equivalents
von: Lee, Jane H., et al.
Veröffentlicht: (2025)
von: Lee, Jane H., et al.
Veröffentlicht: (2025)
Optimized Certainty Equivalent Risk-Controlling Prediction Sets
von: Huang, Jiayi, et al.
Veröffentlicht: (2026)
von: Huang, Jiayi, et al.
Veröffentlicht: (2026)
Howard's Policy Iteration is Subexponential for Deterministic Markov Decision Problems with Rewards of Fixed Bit-size and Arbitrary Discount Factor
von: Mukherjee, Dibyangshu, et al.
Veröffentlicht: (2025)
von: Mukherjee, Dibyangshu, et al.
Veröffentlicht: (2025)
Efficient Computation of Blackwell Optimal Policies using Rational Functions
von: Mukherjee, Dibyangshu, et al.
Veröffentlicht: (2025)
von: Mukherjee, Dibyangshu, et al.
Veröffentlicht: (2025)
A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents
von: Mortensen, Oliver, et al.
Veröffentlicht: (2026)
von: Mortensen, Oliver, et al.
Veröffentlicht: (2026)
Is Transductive Learning Equivalent to PAC Learning?
von: Dughmi, Shaddin, et al.
Veröffentlicht: (2024)
von: Dughmi, Shaddin, et al.
Veröffentlicht: (2024)
A PAC-Bayesian View of Generalisation for Physics-Informed Machine Learning
von: Nguyen, Thien V., et al.
Veröffentlicht: (2026)
von: Nguyen, Thien V., et al.
Veröffentlicht: (2026)
Beyond Non-Degeneracy: Revisiting Certainty Equivalent Heuristic for Online Linear Programming
von: Chen, Yilun, et al.
Veröffentlicht: (2025)
von: Chen, Yilun, et al.
Veröffentlicht: (2025)
The Exact Equivalence of Distance and Kernel Methods for Hypothesis Testing
von: Shen, Cencheng, et al.
Veröffentlicht: (2018)
von: Shen, Cencheng, et al.
Veröffentlicht: (2018)
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
von: Ravindran, Santhosh Kumar
Veröffentlicht: (2025)
von: Ravindran, Santhosh Kumar
Veröffentlicht: (2025)
Distilling the Unknown to Unveil Certainty
von: Zhao, Zhilin, et al.
Veröffentlicht: (2023)
von: Zhao, Zhilin, et al.
Veröffentlicht: (2023)
Calibrating Expressions of Certainty
von: Wang, Peiqi, et al.
Veröffentlicht: (2024)
von: Wang, Peiqi, et al.
Veröffentlicht: (2024)
Time-uniform conformal and PAC prediction
von: Scharfstein, Kayla E., et al.
Veröffentlicht: (2026)
von: Scharfstein, Kayla E., et al.
Veröffentlicht: (2026)
Spectral Bellman Method: Unifying Representation and Exploration in RL
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
Combining Reconstruction and Contrastive Methods for Multimodal Representations in RL
von: Becker, Philipp, et al.
Veröffentlicht: (2023)
von: Becker, Philipp, et al.
Veröffentlicht: (2023)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
Explaining RL Decisions with Trajectories
von: Deshmukh, Shripad Vilasrao, et al.
Veröffentlicht: (2023)
von: Deshmukh, Shripad Vilasrao, et al.
Veröffentlicht: (2023)
On the Design of Safe Continual RL Methods for Control of Nonlinear Systems
von: Coursey, Austin, et al.
Veröffentlicht: (2025)
von: Coursey, Austin, et al.
Veröffentlicht: (2025)
A Framework for Bounding Deterministic Risk with PAC-Bayes: Applications to Majority Votes
von: Leblanc, Benjamin, et al.
Veröffentlicht: (2025)
von: Leblanc, Benjamin, et al.
Veröffentlicht: (2025)
How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance
von: Huang, Jerry Y., et al.
Veröffentlicht: (2026)
von: Huang, Jerry Y., et al.
Veröffentlicht: (2026)
Explaining the Success of Nearest Neighbor Methods in Prediction
von: Chen, George H., et al.
Veröffentlicht: (2025)
von: Chen, George H., et al.
Veröffentlicht: (2025)
A Benchmark Study of Deep-RL Methods for Maximum Coverage Problems over Graphs
von: Liang, Zhicheng, et al.
Veröffentlicht: (2024)
von: Liang, Zhicheng, et al.
Veröffentlicht: (2024)
Embedding Safety into RL: A New Take on Trust Region Methods
von: Milosevic, Nikola, et al.
Veröffentlicht: (2024)
von: Milosevic, Nikola, et al.
Veröffentlicht: (2024)
Know your Trajectory -- Trustworthy Reinforcement Learning deployment through Importance-Based Trajectory Analysis
von: F, Clifford, et al.
Veröffentlicht: (2025)
von: F, Clifford, et al.
Veröffentlicht: (2025)
Multi-View Majority Vote Learning Algorithms: Direct Minimization of PAC-Bayesian Bounds
von: Hennequin, Mehdi, et al.
Veröffentlicht: (2024)
von: Hennequin, Mehdi, et al.
Veröffentlicht: (2024)
AI-powered self-healing enterprise applications: A new era of autonomous systems
von: Guguloth, Praveen Kumar
Veröffentlicht: (2025)
von: Guguloth, Praveen Kumar
Veröffentlicht: (2025)
Gromov-Wasserstein Methods for Multi-View Relational Embedding and Clustering
von: Eufrazio, Rafael Pereira, et al.
Veröffentlicht: (2026)
von: Eufrazio, Rafael Pereira, et al.
Veröffentlicht: (2026)
Bad Habits: Policy Confounding and Out-of-Trajectory Generalization in RL
von: Suau, Miguel, et al.
Veröffentlicht: (2023)
von: Suau, Miguel, et al.
Veröffentlicht: (2023)
Slug Mobile: Test-Bench for RL Testing
von: Morris, Jonathan Wellington, et al.
Veröffentlicht: (2024)
von: Morris, Jonathan Wellington, et al.
Veröffentlicht: (2024)
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
von: Sikchi, Harshit, et al.
Veröffentlicht: (2023)
von: Sikchi, Harshit, et al.
Veröffentlicht: (2023)
Efficient Optimal PAC Learning
von: Høgsgaard, Mikael Møller
Veröffentlicht: (2025)
von: Høgsgaard, Mikael Møller
Veröffentlicht: (2025)
Symmetries in PAC-Bayesian Learning
von: Beck, Armin, et al.
Veröffentlicht: (2025)
von: Beck, Armin, et al.
Veröffentlicht: (2025)
On the Computability of Multiclass PAC Learning
von: Gourdeau, Pascale, et al.
Veröffentlicht: (2025)
von: Gourdeau, Pascale, et al.
Veröffentlicht: (2025)
PAC Learnability in the Presence of Performativity
von: Kirev, Ivan, et al.
Veröffentlicht: (2025)
von: Kirev, Ivan, et al.
Veröffentlicht: (2025)
On the Computability of Robust PAC Learning
von: Gourdeau, Pascale, et al.
Veröffentlicht: (2024)
von: Gourdeau, Pascale, et al.
Veröffentlicht: (2024)
Deep Exploration with PAC-Bayes
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2024)
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
von: Shah, Anvay, et al.
Veröffentlicht: (2026) -
Using Common Random Numbers for Simulation-based Planning with Rollouts
von: Yadav, Sandarbh, et al.
Veröffentlicht: (2026) -
Concentration Bounds for Optimized Certainty Equivalent Risk Estimation
von: Ghosh, Ayon, et al.
Veröffentlicht: (2024) -
Risk-Averse Constrained Reinforcement Learning with Optimized Certainty Equivalents
von: Lee, Jane H., et al.
Veröffentlicht: (2025) -
Optimized Certainty Equivalent Risk-Controlling Prediction Sets
von: Huang, Jiayi, et al.
Veröffentlicht: (2026)