Towards Causal Model-Based Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Caron, Alberto, Mavroudis, Vasilios, Hicks, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Efficient Bayesian Exploration in Model-Based Reinforcement Learning
by: Caron, Alberto, et al.
Published: (2025)
by: Caron, Alberto, et al.
Published: (2025)
A View on Out-of-Distribution Identification from a Statistical Testing Theory Perspective
by: Caron, Alberto, et al.
Published: (2024)
by: Caron, Alberto, et al.
Published: (2024)
Inherently Interpretable and Uncertainty-Aware Models for Online Learning in Cyber-Security Problems
by: Kolicic, Benjamin, et al.
Published: (2024)
by: Kolicic, Benjamin, et al.
Published: (2024)
Entity-based Reinforcement Learning for Autonomous Cyber Defence
by: Thompson, Isaac Symes, et al.
Published: (2024)
by: Thompson, Isaac Symes, et al.
Published: (2024)
Nearest Neighbour with Bandit Feedback
by: Pasteris, Stephen, et al.
Published: (2023)
by: Pasteris, Stephen, et al.
Published: (2023)
Fairness with Exponential Weights
by: Pasteris, Stephen, et al.
Published: (2024)
by: Pasteris, Stephen, et al.
Published: (2024)
Extraction Propagation
by: Pasteris, Stephen, et al.
Published: (2024)
by: Pasteris, Stephen, et al.
Published: (2024)
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
by: Vyas, Sanyam, et al.
Published: (2025)
by: Vyas, Sanyam, et al.
Published: (2025)
Beyond Rewards in Reinforcement Learning for Cyber Defence
by: Bates, Elizabeth, et al.
Published: (2026)
by: Bates, Elizabeth, et al.
Published: (2026)
Less is more? Rewards in RL for Cyber Defence
by: Bates, Elizabeth, et al.
Published: (2025)
by: Bates, Elizabeth, et al.
Published: (2025)
Mitigating Deep Reinforcement Learning Backdoors in the Neural Activation Space
by: Vyas, Sanyam, et al.
Published: (2024)
by: Vyas, Sanyam, et al.
Published: (2024)
Online Convex Optimisation: The Optimal Switching Regret for all Segmentations Simultaneously
by: Pasteris, Stephen, et al.
Published: (2024)
by: Pasteris, Stephen, et al.
Published: (2024)
Autonomous Network Defence using Reinforcement Learning
by: Foley, Myles, et al.
Published: (2024)
by: Foley, Myles, et al.
Published: (2024)
An Attentive Graph Agent for Topology-Adaptive Cyber Defence
by: Sandoval, Ilya Orson, et al.
Published: (2025)
by: Sandoval, Ilya Orson, et al.
Published: (2025)
DRMD: Deep Reinforcement Learning for Malware Detection under Concept Drift
by: McFadden, Shae, et al.
Published: (2025)
by: McFadden, Shae, et al.
Published: (2025)
SoK: The Pitfalls of Deep Reinforcement Learning for Cybersecurity
by: McFadden, Shae, et al.
Published: (2026)
by: McFadden, Shae, et al.
Published: (2026)
Environment Complexity and Nash Equilibria in a Sequential Social Dilemma
by: Yasir, Mustafa, et al.
Published: (2024)
by: Yasir, Mustafa, et al.
Published: (2024)
CybORG++: An Enhanced Gym for the Development of Autonomous Cyber Agents
by: Emerson, Harry, et al.
Published: (2024)
by: Emerson, Harry, et al.
Published: (2024)
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
by: Souly, Alexandra, et al.
Published: (2025)
by: Souly, Alexandra, et al.
Published: (2025)
Zero-Trust Network Access (ZTNA)
by: Mavroudis, Vasilios
Published: (2024)
by: Mavroudis, Vasilios
Published: (2024)
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning
by: Wang, Jingyao, et al.
Published: (2026)
by: Wang, Jingyao, et al.
Published: (2026)
What if we could hot swap our Biometrics?
by: Crowcroft, Jon, et al.
Published: (2025)
by: Crowcroft, Jon, et al.
Published: (2025)
Towards Representation Learning for Weighting Problems in Design-Based Causal Inference
by: Clivio, Oscar, et al.
Published: (2024)
by: Clivio, Oscar, et al.
Published: (2024)
Group Causal Policy Optimization for Post-Training Large Language Models
by: Gu, Ziyin, et al.
Published: (2025)
by: Gu, Ziyin, et al.
Published: (2025)
Quantifying Mix Network Privacy Erosion with Generative Models
by: Mavroudis, Vasilios, et al.
Published: (2025)
by: Mavroudis, Vasilios, et al.
Published: (2025)
STGCN-LSTM for Olympic Medal Prediction: Dynamic Power Modeling and Causal Policy Optimization
by: Wang, Yiquan, et al.
Published: (2025)
by: Wang, Yiquan, et al.
Published: (2025)
Causally-Enhanced Reinforcement Policy Optimization
by: Wang, Xiangqi, et al.
Published: (2025)
by: Wang, Xiangqi, et al.
Published: (2025)
Analysis of Publicly Accessible Operational Technology and Associated Risks
by: Rodda, Matthew, et al.
Published: (2025)
by: Rodda, Matthew, et al.
Published: (2025)
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
by: Ristea, Dan, et al.
Published: (2024)
by: Ristea, Dan, et al.
Published: (2024)
Referential Security as a New Paradigm for AI Evaluations
by: Ristea, Dan, et al.
Published: (2026)
by: Ristea, Dan, et al.
Published: (2026)
Double Horizon Model-Based Policy Optimization
by: Kubo, Akihiro, et al.
Published: (2025)
by: Kubo, Akihiro, et al.
Published: (2025)
Optimal Mixed Integer Linear Optimization Trained Multivariate Classification Trees
by: Alston, Brandon, et al.
Published: (2024)
by: Alston, Brandon, et al.
Published: (2024)
One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image
by: Shereen, Ezzeldin, et al.
Published: (2025)
by: Shereen, Ezzeldin, et al.
Published: (2025)
Towards Unsupervised Causal Representation Learning via Latent Additive Noise Model Causal Autoencoders
by: Ong, Hans Jarett J., et al.
Published: (2025)
by: Ong, Hans Jarett J., et al.
Published: (2025)
Towards Empowerment Gain through Causal Structure Learning in Model-Based RL
by: Cao, Hongye, et al.
Published: (2025)
by: Cao, Hongye, et al.
Published: (2025)
Toward Falsifying Causal Graphs Using a Permutation-Based Test
by: Eulig, Elias, et al.
Published: (2023)
by: Eulig, Elias, et al.
Published: (2023)
Bayesian Sensitivity of Causal Inference Estimators under Evidence-Based Priors
by: Dhawan, Nikita, et al.
Published: (2026)
by: Dhawan, Nikita, et al.
Published: (2026)
Beyond Uniform Credit: Causal Credit Assignment for Policy Optimization
by: Khandoga, Mykola, et al.
Published: (2026)
by: Khandoga, Mykola, et al.
Published: (2026)
Model-Based Reinforcement Learning with Double Oracle Efficiency in Policy Optimization and Offline Estimation
by: Hu, Haichen, et al.
Published: (2026)
by: Hu, Haichen, et al.
Published: (2026)
Learning When to Switch: Adaptive Policy Selection via Reinforcement Learning
by: Tava, Chris
Published: (2025)
by: Tava, Chris
Published: (2025)
Similar Items
-
On Efficient Bayesian Exploration in Model-Based Reinforcement Learning
by: Caron, Alberto, et al.
Published: (2025) -
A View on Out-of-Distribution Identification from a Statistical Testing Theory Perspective
by: Caron, Alberto, et al.
Published: (2024) -
Inherently Interpretable and Uncertainty-Aware Models for Online Learning in Cyber-Security Problems
by: Kolicic, Benjamin, et al.
Published: (2024) -
Entity-based Reinforcement Learning for Autonomous Cyber Defence
by: Thompson, Isaac Symes, et al.
Published: (2024) -
Nearest Neighbour with Bandit Feedback
by: Pasteris, Stephen, et al.
Published: (2023)