Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-Critic
Fuente:
arXiv
Saved in:
| Main Authors: | Ji, Tianying, Luo, Yu, Sun, Fuchun, Zhan, Xianyuan, Zhang, Jianwei, Xu, Huazhe |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
by: Luo, Yu, et al.
Published: (2024)
by: Luo, Yu, et al.
Published: (2024)
ACE : Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
by: Ji, Tianying, et al.
Published: (2024)
by: Ji, Tianying, et al.
Published: (2024)
OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
by: Luo, Yu, et al.
Published: (2024)
by: Luo, Yu, et al.
Published: (2024)
Bidirectional-Reachable Hierarchical Reinforcement Learning with Mutually Responsive Policies
by: Luo, Yu, et al.
Published: (2024)
by: Luo, Yu, et al.
Published: (2024)
RoboGolf: Mastering Real-World Minigolf with a Reflective Multi-Modality Vision-Language Model
by: Zhou, Hantao, et al.
Published: (2024)
by: Zhou, Hantao, et al.
Published: (2024)
Enhancing Multi-Agent Collaboration with Attention-Based Actor-Critic Policies
by: Belinchon, Hugo Garrido-Lestache, et al.
Published: (2025)
by: Belinchon, Hugo Garrido-Lestache, et al.
Published: (2025)
Soft Actor-Critic with Beta Policy via Implicit Reparameterization Gradients
by: Della Libera, Luca
Published: (2024)
by: Della Libera, Luca
Published: (2024)
Rethinking Mutual Information for Language Conditioned Skill Discovery on Imitation Learning
by: Ju, Zhaoxun, et al.
Published: (2024)
by: Ju, Zhaoxun, et al.
Published: (2024)
Improving Value Estimation Critically Enhances Vanilla Policy Gradient
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration
by: Bai, Qinxun, et al.
Published: (2025)
by: Bai, Qinxun, et al.
Published: (2025)
The Paradox of Success in Evolutionary and Bioinspired Optimization: Revisiting Critical Issues, Key Studies, and Methodological Pathways
by: Molina, Daniel, et al.
Published: (2025)
by: Molina, Daniel, et al.
Published: (2025)
Off-Policy Correction For Multi-Agent Reinforcement Learning
by: Zawalski, Michał, et al.
Published: (2021)
by: Zawalski, Michał, et al.
Published: (2021)
Enabling Off-Policy Imitation Learning with Deep Actor Critic Stabilization
by: Sen, Sayambhu, et al.
Published: (2025)
by: Sen, Sayambhu, et al.
Published: (2025)
Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory
by: Jiang, Rongjie, et al.
Published: (2026)
by: Jiang, Rongjie, et al.
Published: (2026)
Off-Policy Actor-Critic with Sigmoid-Bounded Entropy for Real-World Robot Learning
by: Wu, Xiefeng, et al.
Published: (2026)
by: Wu, Xiefeng, et al.
Published: (2026)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
by: Sun, Kangkang, et al.
Published: (2026)
by: Sun, Kangkang, et al.
Published: (2026)
Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment
by: Li, Bobo, et al.
Published: (2026)
by: Li, Bobo, et al.
Published: (2026)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation
by: Chen, Junting, et al.
Published: (2024)
by: Chen, Junting, et al.
Published: (2024)
Value Improved Actor Critic Algorithms
by: Oren, Yaniv, et al.
Published: (2024)
by: Oren, Yaniv, et al.
Published: (2024)
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
by: Hong, Yoosung
Published: (2026)
by: Hong, Yoosung
Published: (2026)
SALLIE: Safeguarding Against Latent Language & Image Exploits
by: Azov, Guy, et al.
Published: (2026)
by: Azov, Guy, et al.
Published: (2026)
Expressive Value Learning for Scalable Offline Reinforcement Learning
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation
by: Nakanishi, Kosuke, et al.
Published: (2025)
by: Nakanishi, Kosuke, et al.
Published: (2025)
Logging Policy Design for Off-Policy Evaluation
by: Douglas, Connor, et al.
Published: (2026)
by: Douglas, Connor, et al.
Published: (2026)
Re-Mask and Redirect: Exploiting Denoising Irreversibility in Diffusion Language Models
by: Singh, Arth
Published: (2026)
by: Singh, Arth
Published: (2026)
Robust and Diverse Multi-Agent Learning via Rational Policy Gradient
by: Lauffer, Niklas, et al.
Published: (2025)
by: Lauffer, Niklas, et al.
Published: (2025)
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
Partially Observable Reference Policy Programming: Solving POMDPs Sans Numerical Optimisation
by: Kim, Edward, et al.
Published: (2025)
by: Kim, Edward, et al.
Published: (2025)
Exploiting contextual information to improve stance detection in informal political discourse with LLMs
by: Sucu, Arman Engin, et al.
Published: (2026)
by: Sucu, Arman Engin, et al.
Published: (2026)
Evaluation of Black-Box XAI Approaches for Predictors of Values of Boolean Formulae
by: Armoni-Friedmann, Stav, et al.
Published: (2025)
by: Armoni-Friedmann, Stav, et al.
Published: (2025)
Optimization of Activity Batching Policies in Business Processes
by: López-Pintado, Orlenys, et al.
Published: (2025)
by: López-Pintado, Orlenys, et al.
Published: (2025)
Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
by: Liao, Jianxing, et al.
Published: (2025)
by: Liao, Jianxing, et al.
Published: (2025)
Critical Insights into Leading Conversational AI Models
by: Kohli, Urja, et al.
Published: (2025)
by: Kohli, Urja, et al.
Published: (2025)
Exploiting Causality Signals in Medical Images: A Pilot Study with Empirical Results
by: Carloni, Gianluca, et al.
Published: (2023)
by: Carloni, Gianluca, et al.
Published: (2023)
Process Supervision-Guided Policy Optimization for Code Generation
by: Dai, Ning, et al.
Published: (2024)
by: Dai, Ning, et al.
Published: (2024)
Dynamic Policy Induction for Adaptive Prompt Optimization: Bridging the Efficiency-Accuracy Gap via Lightweight Reinforcement Learning
by: Xu, Jiexi
Published: (2025)
by: Xu, Jiexi
Published: (2025)
Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Training
by: Gu, Zhengyao, et al.
Published: (2026)
by: Gu, Zhengyao, et al.
Published: (2026)
GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies
by: Zhang, He, et al.
Published: (2026)
by: Zhang, He, et al.
Published: (2026)
Similar Items
-
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
by: Luo, Yu, et al.
Published: (2024) -
ACE : Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
by: Ji, Tianying, et al.
Published: (2024) -
OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
by: Luo, Yu, et al.
Published: (2024) -
Bidirectional-Reachable Hierarchical Reinforcement Learning with Mutually Responsive Policies
by: Luo, Yu, et al.
Published: (2024) -
RoboGolf: Mastering Real-World Minigolf with a Reflective Multi-Modality Vision-Language Model
by: Zhou, Hantao, et al.
Published: (2024)