A View of the Certainty-Equivalence Method for PAC RL as an Application of the Trajectory Tree Method
Fuente:
arXiv
Guardado en:
| Autores principales: | Kalyanakrishnan, Shivaram, Shah, Sheel, Guguloth, Santhosh Kumar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
por: Shah, Anvay, et al.
Publicado: (2026)
por: Shah, Anvay, et al.
Publicado: (2026)
Using Common Random Numbers for Simulation-based Planning with Rollouts
por: Yadav, Sandarbh, et al.
Publicado: (2026)
por: Yadav, Sandarbh, et al.
Publicado: (2026)
Concentration Bounds for Optimized Certainty Equivalent Risk Estimation
por: Ghosh, Ayon, et al.
Publicado: (2024)
por: Ghosh, Ayon, et al.
Publicado: (2024)
Risk-Averse Constrained Reinforcement Learning with Optimized Certainty Equivalents
por: Lee, Jane H., et al.
Publicado: (2025)
por: Lee, Jane H., et al.
Publicado: (2025)
Optimized Certainty Equivalent Risk-Controlling Prediction Sets
por: Huang, Jiayi, et al.
Publicado: (2026)
por: Huang, Jiayi, et al.
Publicado: (2026)
Howard's Policy Iteration is Subexponential for Deterministic Markov Decision Problems with Rewards of Fixed Bit-size and Arbitrary Discount Factor
por: Mukherjee, Dibyangshu, et al.
Publicado: (2025)
por: Mukherjee, Dibyangshu, et al.
Publicado: (2025)
Efficient Computation of Blackwell Optimal Policies using Rational Functions
por: Mukherjee, Dibyangshu, et al.
Publicado: (2025)
por: Mukherjee, Dibyangshu, et al.
Publicado: (2025)
A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents
por: Wang, Kaiwen, et al.
Publicado: (2024)
por: Wang, Kaiwen, et al.
Publicado: (2024)
On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents
por: Mortensen, Oliver, et al.
Publicado: (2026)
por: Mortensen, Oliver, et al.
Publicado: (2026)
Is Transductive Learning Equivalent to PAC Learning?
por: Dughmi, Shaddin, et al.
Publicado: (2024)
por: Dughmi, Shaddin, et al.
Publicado: (2024)
A PAC-Bayesian View of Generalisation for Physics-Informed Machine Learning
por: Nguyen, Thien V., et al.
Publicado: (2026)
por: Nguyen, Thien V., et al.
Publicado: (2026)
Beyond Non-Degeneracy: Revisiting Certainty Equivalent Heuristic for Online Linear Programming
por: Chen, Yilun, et al.
Publicado: (2025)
por: Chen, Yilun, et al.
Publicado: (2025)
The Exact Equivalence of Distance and Kernel Methods for Hypothesis Testing
por: Shen, Cencheng, et al.
Publicado: (2018)
por: Shen, Cencheng, et al.
Publicado: (2018)
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
por: Ravindran, Santhosh Kumar
Publicado: (2025)
por: Ravindran, Santhosh Kumar
Publicado: (2025)
Distilling the Unknown to Unveil Certainty
por: Zhao, Zhilin, et al.
Publicado: (2023)
por: Zhao, Zhilin, et al.
Publicado: (2023)
Calibrating Expressions of Certainty
por: Wang, Peiqi, et al.
Publicado: (2024)
por: Wang, Peiqi, et al.
Publicado: (2024)
Time-uniform conformal and PAC prediction
por: Scharfstein, Kayla E., et al.
Publicado: (2026)
por: Scharfstein, Kayla E., et al.
Publicado: (2026)
Spectral Bellman Method: Unifying Representation and Exploration in RL
por: Nabati, Ofir, et al.
Publicado: (2025)
por: Nabati, Ofir, et al.
Publicado: (2025)
Combining Reconstruction and Contrastive Methods for Multimodal Representations in RL
por: Becker, Philipp, et al.
Publicado: (2023)
por: Becker, Philipp, et al.
Publicado: (2023)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
por: Chen, Zihan, et al.
Publicado: (2025)
por: Chen, Zihan, et al.
Publicado: (2025)
Explaining RL Decisions with Trajectories
por: Deshmukh, Shripad Vilasrao, et al.
Publicado: (2023)
por: Deshmukh, Shripad Vilasrao, et al.
Publicado: (2023)
On the Design of Safe Continual RL Methods for Control of Nonlinear Systems
por: Coursey, Austin, et al.
Publicado: (2025)
por: Coursey, Austin, et al.
Publicado: (2025)
A Framework for Bounding Deterministic Risk with PAC-Bayes: Applications to Majority Votes
por: Leblanc, Benjamin, et al.
Publicado: (2025)
por: Leblanc, Benjamin, et al.
Publicado: (2025)
How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance
por: Huang, Jerry Y., et al.
Publicado: (2026)
por: Huang, Jerry Y., et al.
Publicado: (2026)
Explaining the Success of Nearest Neighbor Methods in Prediction
por: Chen, George H., et al.
Publicado: (2025)
por: Chen, George H., et al.
Publicado: (2025)
A Benchmark Study of Deep-RL Methods for Maximum Coverage Problems over Graphs
por: Liang, Zhicheng, et al.
Publicado: (2024)
por: Liang, Zhicheng, et al.
Publicado: (2024)
Embedding Safety into RL: A New Take on Trust Region Methods
por: Milosevic, Nikola, et al.
Publicado: (2024)
por: Milosevic, Nikola, et al.
Publicado: (2024)
Know your Trajectory -- Trustworthy Reinforcement Learning deployment through Importance-Based Trajectory Analysis
por: F, Clifford, et al.
Publicado: (2025)
por: F, Clifford, et al.
Publicado: (2025)
Multi-View Majority Vote Learning Algorithms: Direct Minimization of PAC-Bayesian Bounds
por: Hennequin, Mehdi, et al.
Publicado: (2024)
por: Hennequin, Mehdi, et al.
Publicado: (2024)
AI-powered self-healing enterprise applications: A new era of autonomous systems
por: Guguloth, Praveen Kumar
Publicado: (2025)
por: Guguloth, Praveen Kumar
Publicado: (2025)
Gromov-Wasserstein Methods for Multi-View Relational Embedding and Clustering
por: Eufrazio, Rafael Pereira, et al.
Publicado: (2026)
por: Eufrazio, Rafael Pereira, et al.
Publicado: (2026)
Bad Habits: Policy Confounding and Out-of-Trajectory Generalization in RL
por: Suau, Miguel, et al.
Publicado: (2023)
por: Suau, Miguel, et al.
Publicado: (2023)
Slug Mobile: Test-Bench for RL Testing
por: Morris, Jonathan Wellington, et al.
Publicado: (2024)
por: Morris, Jonathan Wellington, et al.
Publicado: (2024)
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
por: Sikchi, Harshit, et al.
Publicado: (2023)
por: Sikchi, Harshit, et al.
Publicado: (2023)
Efficient Optimal PAC Learning
por: Høgsgaard, Mikael Møller
Publicado: (2025)
por: Høgsgaard, Mikael Møller
Publicado: (2025)
Symmetries in PAC-Bayesian Learning
por: Beck, Armin, et al.
Publicado: (2025)
por: Beck, Armin, et al.
Publicado: (2025)
On the Computability of Multiclass PAC Learning
por: Gourdeau, Pascale, et al.
Publicado: (2025)
por: Gourdeau, Pascale, et al.
Publicado: (2025)
PAC Learnability in the Presence of Performativity
por: Kirev, Ivan, et al.
Publicado: (2025)
por: Kirev, Ivan, et al.
Publicado: (2025)
On the Computability of Robust PAC Learning
por: Gourdeau, Pascale, et al.
Publicado: (2024)
por: Gourdeau, Pascale, et al.
Publicado: (2024)
Deep Exploration with PAC-Bayes
por: Tasdighi, Bahareh, et al.
Publicado: (2024)
por: Tasdighi, Bahareh, et al.
Publicado: (2024)
Ejemplares similares
-
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
por: Shah, Anvay, et al.
Publicado: (2026) -
Using Common Random Numbers for Simulation-based Planning with Rollouts
por: Yadav, Sandarbh, et al.
Publicado: (2026) -
Concentration Bounds for Optimized Certainty Equivalent Risk Estimation
por: Ghosh, Ayon, et al.
Publicado: (2024) -
Risk-Averse Constrained Reinforcement Learning with Optimized Certainty Equivalents
por: Lee, Jane H., et al.
Publicado: (2025) -
Optimized Certainty Equivalent Risk-Controlling Prediction Sets
por: Huang, Jiayi, et al.
Publicado: (2026)