Saved in:
| Main Authors: | Cloete, Jacques, Vertovec, Nikolaus, Abate, Alessandro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.06386 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PlatoLTL: Learning to Generalize Across Symbols in LTL Instructions for Multi-Task RL
by: Cloete, Jacques, et al.
Published: (2026)
by: Cloete, Jacques, et al.
Published: (2026)
Certified Neural Approximations of Nonlinear Dynamics
by: Mathiesen, Frederik Baymler, et al.
Published: (2025)
by: Mathiesen, Frederik Baymler, et al.
Published: (2025)
Zero-Shot Instruction Following in RL via Structured LTL Representations
by: Jackermeier, Mathias, et al.
Published: (2026)
by: Jackermeier, Mathias, et al.
Published: (2026)
Certified Approximate Reachability (CARe): Formal Error Bounds on Deep Learning of Reachable Sets
by: Solanki, Prashant, et al.
Published: (2025)
by: Solanki, Prashant, et al.
Published: (2025)
Certifiably Robust Policies for Uncertain Parametric Environments
by: Schnitzer, Yannik, et al.
Published: (2024)
by: Schnitzer, Yannik, et al.
Published: (2024)
Scalable Verification of Neural Control Barrier Functions Using Linear Bound Propagation
by: Vertovec, Nikolaus, et al.
Published: (2025)
by: Vertovec, Nikolaus, et al.
Published: (2025)
Finite sample learning of moving targets
by: Vertovec, Nikolaus, et al.
Published: (2024)
by: Vertovec, Nikolaus, et al.
Published: (2024)
DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RL
by: Jackermeier, Mathias, et al.
Published: (2024)
by: Jackermeier, Mathias, et al.
Published: (2024)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
by: Melo, Luckeciano C., et al.
Published: (2025)
by: Melo, Luckeciano C., et al.
Published: (2025)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)
by: Griffin, Charlie, et al.
Published: (2024)
Zero-Shot Instruction Following in RL via Structured LTL Representations
by: Giuri, Mattia, et al.
Published: (2025)
by: Giuri, Mattia, et al.
Published: (2025)
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
Data-Driven Policy Mapping for Safe RL-based Energy Management Systems
by: Zangato, Theo, et al.
Published: (2025)
by: Zangato, Theo, et al.
Published: (2025)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
SPoT: Subpixel Placement of Tokens in Vision Transformers
by: Hjelkrem-Tan, Martine, et al.
Published: (2025)
by: Hjelkrem-Tan, Martine, et al.
Published: (2025)
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL
by: Chen, Zhaoyang, et al.
Published: (2025)
by: Chen, Zhaoyang, et al.
Published: (2025)
Action-Free Offline-to-Online RL via Discretised State Policies
by: Neggatu, Natinael Solomon, et al.
Published: (2026)
by: Neggatu, Natinael Solomon, et al.
Published: (2026)
SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data
by: Romeo, Carlo, et al.
Published: (2026)
by: Romeo, Carlo, et al.
Published: (2026)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
by: Ankile, Lars, et al.
Published: (2025)
by: Ankile, Lars, et al.
Published: (2025)
JaxRobotarium: Training and Deploying Multi-Robot Policies in 10 Minutes
by: Jain, Shalin Anand, et al.
Published: (2025)
by: Jain, Shalin Anand, et al.
Published: (2025)
Probabilistic Performance Guarantees for Multi-Task Reinforcement Learning
by: Schnitzer, Yannik, et al.
Published: (2026)
by: Schnitzer, Yannik, et al.
Published: (2026)
Policy Bifurcation in Safe Reinforcement Learning
by: Zou, Wenjun, et al.
Published: (2024)
by: Zou, Wenjun, et al.
Published: (2024)
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
On-Policy RL with Optimal Reward Baseline
by: Hao, Yaru, et al.
Published: (2025)
by: Hao, Yaru, et al.
Published: (2025)
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026)
by: Mathur, Puneet, et al.
Published: (2026)
Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL
by: Sandhu, Dillon, et al.
Published: (2026)
by: Sandhu, Dillon, et al.
Published: (2026)
Safe Deep Policy Adaptation
by: Xiao, Wenli, et al.
Published: (2023)
by: Xiao, Wenli, et al.
Published: (2023)
Neural Proofs for Sound Verification and Control of Complex Systems
by: Abate, Alessandro
Published: (2025)
by: Abate, Alessandro
Published: (2025)
Inference Time Policy Optimization for Offline RL with Differentiable World Models
by: Deb, Rohan, et al.
Published: (2026)
by: Deb, Rohan, et al.
Published: (2026)
Decision-Point Guided Safe Policy Improvement
by: Sharma, Abhishek, et al.
Published: (2024)
by: Sharma, Abhishek, et al.
Published: (2024)
Training-Free Imitation Learning with Closed-Form Diffusion Policies
by: Mishra, Raghav, et al.
Published: (2026)
by: Mishra, Raghav, et al.
Published: (2026)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
SMaRt: Improving GANs with Score Matching Regularity
by: Xia, Mengfei, et al.
Published: (2023)
by: Xia, Mengfei, et al.
Published: (2023)
SNPL: Simultaneous Policy Learning and Evaluation for Safe Multi-Objective Policy Improvement
by: Cho, Brian, et al.
Published: (2025)
by: Cho, Brian, et al.
Published: (2025)
CSPI-MT: Calibrated Safe Policy Improvement with Multiple Testing for Threshold Policies
by: Cho, Brian M, et al.
Published: (2024)
by: Cho, Brian M, et al.
Published: (2024)
RL in the Wild: Characterizing RLVR Training in LLM Deployment
by: Zhou, Jiecheng, et al.
Published: (2025)
by: Zhou, Jiecheng, et al.
Published: (2025)
Similar Items
-
PlatoLTL: Learning to Generalize Across Symbols in LTL Instructions for Multi-Task RL
by: Cloete, Jacques, et al.
Published: (2026) -
Certified Neural Approximations of Nonlinear Dynamics
by: Mathiesen, Frederik Baymler, et al.
Published: (2025) -
Zero-Shot Instruction Following in RL via Structured LTL Representations
by: Jackermeier, Mathias, et al.
Published: (2026) -
Certified Approximate Reachability (CARe): Formal Error Bounds on Deep Learning of Reachable Sets
by: Solanki, Prashant, et al.
Published: (2025) -
Certifiably Robust Policies for Uncertain Parametric Environments
by: Schnitzer, Yannik, et al.
Published: (2024)