Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk
Fuente:
arXiv
Saved in:
| Main Authors: | Woydt, Tim, Zuercher, Paul-David |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fast Rates for Bandit PAC Multiclass Classification
by: Erez, Liad, et al.
Published: (2024)
by: Erez, Liad, et al.
Published: (2024)
Constraint-Anchored Attribution: Feasibility-Certified Counterfactuals and Bonferroni-PAC Sufficient Subsets for Neural CO Policies
by: Lafifi, Sohaib
Published: (2026)
by: Lafifi, Sohaib
Published: (2026)
SPO: Sequential Monte Carlo Policy Optimisation
by: Macfarlane, Matthew V, et al.
Published: (2024)
by: Macfarlane, Matthew V, et al.
Published: (2024)
Holder Policy Optimisation
by: Chen, Yuxiang, et al.
Published: (2026)
by: Chen, Yuxiang, et al.
Published: (2026)
PAC-Bayesian Reinforcement Learning Trains Generalizable Policies
by: Zitouni, Abdelkrim, et al.
Published: (2025)
by: Zitouni, Abdelkrim, et al.
Published: (2025)
Causal Contextual Bandits with Adaptive Context
by: Madhavan, Rahul, et al.
Published: (2024)
by: Madhavan, Rahul, et al.
Published: (2024)
Causally Abstracted Multi-armed Bandits
by: Zennaro, Fabio Massimo, et al.
Published: (2024)
by: Zennaro, Fabio Massimo, et al.
Published: (2024)
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Fodor and Pylyshyn's Legacy: Still No Human-like Systematic Compositionality in Neural Networks
by: Woydt, Tim, et al.
Published: (2025)
by: Woydt, Tim, et al.
Published: (2025)
The Minimal Search Space for Conditional Causal Bandits
by: Simoes, Francisco N. F. Q., et al.
Published: (2025)
by: Simoes, Francisco N. F. Q., et al.
Published: (2025)
Certifiably Robust Policies for Uncertain Parametric Environments
by: Schnitzer, Yannik, et al.
Published: (2024)
by: Schnitzer, Yannik, et al.
Published: (2024)
Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?
by: Wüst, Antonia, et al.
Published: (2024)
by: Wüst, Antonia, et al.
Published: (2024)
Risk-aware Direct Preference Optimization under Nested Risk Measure
by: Zhang, Lijun, et al.
Published: (2025)
by: Zhang, Lijun, et al.
Published: (2025)
Mirror Learning: A Unifying Framework of Policy Optimisation
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees
by: Zhang, Zuyuan, et al.
Published: (2025)
by: Zhang, Zuyuan, et al.
Published: (2025)
HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees
by: Zeng, Hao, et al.
Published: (2026)
by: Zeng, Hao, et al.
Published: (2026)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
by: Hou, Yunlong, et al.
Published: (2025)
by: Hou, Yunlong, et al.
Published: (2025)
ECPO: Evidence-Coupled Policy Optimization for Evidence-Certified Candidate Ranking
by: Hu, Miaobo, et al.
Published: (2026)
by: Hu, Miaobo, et al.
Published: (2026)
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
by: Shah, Anvay, et al.
Published: (2026)
by: Shah, Anvay, et al.
Published: (2026)
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
by: Shimizu, Tatsuhiro, et al.
Published: (2024)
by: Shimizu, Tatsuhiro, et al.
Published: (2024)
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation
by: Mead, Harry, et al.
Published: (2025)
by: Mead, Harry, et al.
Published: (2025)
Generalization Bounds: Perspectives from Information Theory and PAC-Bayes
by: Hellström, Fredrik, et al.
Published: (2023)
by: Hellström, Fredrik, et al.
Published: (2023)
PAC Privacy Preserving Diffusion Models
by: Xu, Qipan, et al.
Published: (2023)
by: Xu, Qipan, et al.
Published: (2023)
Turning Sand to Gold: Recycling Data to Bridge On-Policy and Off-Policy Learning via Causal Bound
by: Fiskus, Tal, et al.
Published: (2025)
by: Fiskus, Tal, et al.
Published: (2025)
CausalGDP: Causality-Guided Diffusion Policies for Reinforcement Learning
by: Xiao, Xiaofeng, et al.
Published: (2026)
by: Xiao, Xiaofeng, et al.
Published: (2026)
Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards
by: Liaw, Sarah, et al.
Published: (2025)
by: Liaw, Sarah, et al.
Published: (2025)
How Catastrophic is Your LLM? Certifying Risk in Conversation
by: Wang, Chengxiao, et al.
Published: (2025)
by: Wang, Chengxiao, et al.
Published: (2025)
Root Cause Attribution of Delivery Risks via Causal Discovery with Reinforcement Learning
by: Xiao, Minheng
Published: (2024)
by: Xiao, Minheng
Published: (2024)
Certified Adversarial Robustness via Partition-based Randomized Smoothing
by: Goli, Hossein, et al.
Published: (2024)
by: Goli, Hossein, et al.
Published: (2024)
Optimizing Warfarin Dosing Using Contextual Bandit: An Offline Policy Learning and Evaluation Method
by: Huang, Yong, et al.
Published: (2024)
by: Huang, Yong, et al.
Published: (2024)
Is Efficient PAC Learning Possible with an Oracle That Responds 'Yes' or 'No'?
by: Daskalakis, Constantinos, et al.
Published: (2024)
by: Daskalakis, Constantinos, et al.
Published: (2024)
Learning by Doing: An Online Causal Reinforcement Learning Framework with Causal-Aware Policy
by: Cai, Ruichu, et al.
Published: (2024)
by: Cai, Ruichu, et al.
Published: (2024)
QuantFPFlow: Quantum Amplitude Estimation for Fokker--Planck Policy Optimisation in Continuous Reinforcement Learning
by: Weinberg, Abraham Itzhak
Published: (2026)
by: Weinberg, Abraham Itzhak
Published: (2026)
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
by: Ren, Tao, et al.
Published: (2025)
by: Ren, Tao, et al.
Published: (2025)
Refined PAC-Bayes Bounds for Offline Bandits
by: Gouverneur, Amaury, et al.
Published: (2025)
by: Gouverneur, Amaury, et al.
Published: (2025)
A Diffusion Analysis of Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)
by: Lattimore, Tor
Published: (2026)
Certified Robustness against Sparse Adversarial Perturbations via Data Localization
by: Pal, Ambar, et al.
Published: (2024)
by: Pal, Ambar, et al.
Published: (2024)
Certified Robustness via Dynamic Margin Maximization and Improved Lipschitz Regularization
by: Fazlyab, Mahyar, et al.
Published: (2023)
by: Fazlyab, Mahyar, et al.
Published: (2023)
Certified Robustness for Deep Equilibrium Models via Serialized Random Smoothing
by: Gao, Weizhi, et al.
Published: (2024)
by: Gao, Weizhi, et al.
Published: (2024)
Neural Sum-of-Squares: Certifying the Nonnegativity of Polynomials with Transformers
by: Pelleriti, Nico, et al.
Published: (2025)
by: Pelleriti, Nico, et al.
Published: (2025)
Similar Items
-
Fast Rates for Bandit PAC Multiclass Classification
by: Erez, Liad, et al.
Published: (2024) -
Constraint-Anchored Attribution: Feasibility-Certified Counterfactuals and Bonferroni-PAC Sufficient Subsets for Neural CO Policies
by: Lafifi, Sohaib
Published: (2026) -
SPO: Sequential Monte Carlo Policy Optimisation
by: Macfarlane, Matthew V, et al.
Published: (2024) -
Holder Policy Optimisation
by: Chen, Yuxiang, et al.
Published: (2026) -
PAC-Bayesian Reinforcement Learning Trains Generalizable Policies
by: Zitouni, Abdelkrim, et al.
Published: (2025)