A Case for Validation Buffer in Pessimistic Actor-Critic
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nauman, Michal, Ostaszewski, Mateusz, Cygan, Marek |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
von: Nauman, Michal, et al.
Veröffentlicht: (2024)
von: Nauman, Michal, et al.
Veröffentlicht: (2024)
On the Theory of Risk-Aware Agents: Bridging Actor-Critic and Economics
von: Nauman, Michal, et al.
Veröffentlicht: (2023)
von: Nauman, Michal, et al.
Veröffentlicht: (2023)
Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
von: Nauman, Michal, et al.
Veröffentlicht: (2024)
von: Nauman, Michal, et al.
Veröffentlicht: (2024)
Reward-Conditioned Reinforcement Learning
von: Nauman, Michal, et al.
Veröffentlicht: (2026)
von: Nauman, Michal, et al.
Veröffentlicht: (2026)
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
von: Nauman, Michal, et al.
Veröffentlicht: (2025)
von: Nauman, Michal, et al.
Veröffentlicht: (2025)
Is Temporal Difference Learning the Gold Standard for Stitching in RL?
von: Bortkiewicz, Michał, et al.
Veröffentlicht: (2025)
von: Bortkiewicz, Michał, et al.
Veröffentlicht: (2025)
Balancing Expressivity and Robustness: Constrained Rational Activations for Reinforcement Learning
von: Surdej, Rafał, et al.
Veröffentlicht: (2025)
von: Surdej, Rafał, et al.
Veröffentlicht: (2025)
Debate2Create: Robot Co-design via Multi-Agent LLM Debate
von: Qiu, Kevin, et al.
Veröffentlicht: (2025)
von: Qiu, Kevin, et al.
Veröffentlicht: (2025)
Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
von: Wołczyk, Maciej, et al.
Veröffentlicht: (2024)
von: Wołczyk, Maciej, et al.
Veröffentlicht: (2024)
Actor-Critic without Actor
von: Ki, Donghyeon, et al.
Veröffentlicht: (2025)
von: Ki, Donghyeon, et al.
Veröffentlicht: (2025)
Pessimistic Off-Policy Optimization for Learning to Rank
von: Cief, Matej, et al.
Veröffentlicht: (2022)
von: Cief, Matej, et al.
Veröffentlicht: (2022)
Actor-Critic or Critic-Actor? A Tale of Two Time Scales
von: Bhatnagar, Shalabh, et al.
Veröffentlicht: (2022)
von: Bhatnagar, Shalabh, et al.
Veröffentlicht: (2022)
Learning Continually by Spectral Regularization
von: Lewandowski, Alex, et al.
Veröffentlicht: (2024)
von: Lewandowski, Alex, et al.
Veröffentlicht: (2024)
FlySearch: Exploring how vision-language models explore
von: Pardyl, Adam, et al.
Veröffentlicht: (2025)
von: Pardyl, Adam, et al.
Veröffentlicht: (2025)
floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
von: Agrawalla, Bhavya, et al.
Veröffentlicht: (2025)
von: Agrawalla, Bhavya, et al.
Veröffentlicht: (2025)
What Matters in Hierarchical Search for Combinatorial Reasoning Problems?
von: Zawalski, Michał, et al.
Veröffentlicht: (2024)
von: Zawalski, Michał, et al.
Veröffentlicht: (2024)
Vid2Sid: Videos Can Help Close the Sim2Real Gap
von: Qiu, Kevin, et al.
Veröffentlicht: (2026)
von: Qiu, Kevin, et al.
Veröffentlicht: (2026)
D2 Actor Critic: Diffusion Actor Meets Distributional Critic
von: Zhang, Lunjun, et al.
Veröffentlicht: (2025)
von: Zhang, Lunjun, et al.
Veröffentlicht: (2025)
Actor-Critic Reinforcement Learning with Phased Actor
von: Wu, Ruofan, et al.
Veröffentlicht: (2024)
von: Wu, Ruofan, et al.
Veröffentlicht: (2024)
Generative Actor Critic
von: Qin, Aoyang, et al.
Veröffentlicht: (2025)
von: Qin, Aoyang, et al.
Veröffentlicht: (2025)
$μ$-Parametrization for Mixture of Experts
von: Małaśnicki, Jan, et al.
Veröffentlicht: (2025)
von: Małaśnicki, Jan, et al.
Veröffentlicht: (2025)
Pessimistic Backward Policy for GFlowNets
von: Jang, Hyosoon, et al.
Veröffentlicht: (2024)
von: Jang, Hyosoon, et al.
Veröffentlicht: (2024)
PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
von: Chen, Yimeng, et al.
Veröffentlicht: (2025)
von: Chen, Yimeng, et al.
Veröffentlicht: (2025)
Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization
von: Masarczyk, Wojciech, et al.
Veröffentlicht: (2025)
von: Masarczyk, Wojciech, et al.
Veröffentlicht: (2025)
Finite-Time Analysis of Three-Timescale Constrained Actor-Critic and Constrained Natural Actor-Critic Algorithms
von: Panda, Prashansa, et al.
Veröffentlicht: (2023)
von: Panda, Prashansa, et al.
Veröffentlicht: (2023)
Hyperparameter Tuning Through Pessimistic Bilevel Optimization
von: Ustun, Meltem Apaydin, et al.
Veröffentlicht: (2024)
von: Ustun, Meltem Apaydin, et al.
Veröffentlicht: (2024)
Learning a Pessimistic Reward Model in RLHF
von: Xu, Yinglun, et al.
Veröffentlicht: (2025)
von: Xu, Yinglun, et al.
Veröffentlicht: (2025)
On the Convergence of Single-Timescale Actor-Critic
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
Adaptive Ensemble Aggregation for Actor-Critics
von: Werge, Nicklas, et al.
Veröffentlicht: (2025)
von: Werge, Nicklas, et al.
Veröffentlicht: (2025)
Risk-Sensitive Exponential Actor Critic
von: Granados, Alonso, et al.
Veröffentlicht: (2026)
von: Granados, Alonso, et al.
Veröffentlicht: (2026)
Safe Langevin Soft Actor Critic
von: Keswani, Mahesh, et al.
Veröffentlicht: (2026)
von: Keswani, Mahesh, et al.
Veröffentlicht: (2026)
Actor-Critic with Active Importance Sampling
von: Molaei, Majid, et al.
Veröffentlicht: (2026)
von: Molaei, Majid, et al.
Veröffentlicht: (2026)
What Does Flow Matching Bring To TD Learning?
von: Agrawalla, Bhavya, et al.
Veröffentlicht: (2026)
von: Agrawalla, Bhavya, et al.
Veröffentlicht: (2026)
A Theoretical Justification for Asymmetric Actor-Critic Algorithms
von: Lambrechts, Gaspard, et al.
Veröffentlicht: (2025)
von: Lambrechts, Gaspard, et al.
Veröffentlicht: (2025)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024)
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024)
Pessimistic Risk-Aware Policy Learning in Contextual Bandits
von: Wan, Yilong, et al.
Veröffentlicht: (2026)
von: Wan, Yilong, et al.
Veröffentlicht: (2026)
Compatible Gradient Approximations for Actor-Critic Algorithms
von: Saglam, Baturay, et al.
Veröffentlicht: (2024)
von: Saglam, Baturay, et al.
Veröffentlicht: (2024)
Refined Analysis of Entropy-Regularized Actor-Critic
von: Labbi, Safwan, et al.
Veröffentlicht: (2026)
von: Labbi, Safwan, et al.
Veröffentlicht: (2026)
Actor-Critic Pretraining for Proximal Policy Optimization
von: Kernbach, Andreas, et al.
Veröffentlicht: (2026)
von: Kernbach, Andreas, et al.
Veröffentlicht: (2026)
Deep Actor-Critics with Tight Risk Certificates
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2025)
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
von: Nauman, Michal, et al.
Veröffentlicht: (2024) -
On the Theory of Risk-Aware Agents: Bridging Actor-Critic and Economics
von: Nauman, Michal, et al.
Veröffentlicht: (2023) -
Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
von: Nauman, Michal, et al.
Veröffentlicht: (2024) -
Reward-Conditioned Reinforcement Learning
von: Nauman, Michal, et al.
Veröffentlicht: (2026) -
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
von: Nauman, Michal, et al.
Veröffentlicht: (2025)