Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Poole, Benjamin, Quinn, Andrew, Yang, Li, Lee, Minwoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Interactive Reinforcement Learning with Intrinsic Feedback
von: Poole, Benjamin, et al.
Veröffentlicht: (2021)
von: Poole, Benjamin, et al.
Veröffentlicht: (2021)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
von: Khoriaty, Matthew, et al.
Veröffentlicht: (2025)
von: Khoriaty, Matthew, et al.
Veröffentlicht: (2025)
Don't Forget Imagination!
von: Vityaev, Evgenii E., et al.
Veröffentlicht: (2025)
von: Vityaev, Evgenii E., et al.
Veröffentlicht: (2025)
Continual Diffuser (CoD): Mastering Continual Offline Reinforcement Learning with Experience Rehearsal
von: Hu, Jifeng, et al.
Veröffentlicht: (2024)
von: Hu, Jifeng, et al.
Veröffentlicht: (2024)
Trust, Don't Trust, or Flip: Robust Preference-Based Reinforcement Learning with Multi-Expert Feedback
von: Hosseini, Seyed Amir, et al.
Veröffentlicht: (2026)
von: Hosseini, Seyed Amir, et al.
Veröffentlicht: (2026)
Transformers Don't In-Context Learn Least Squares Regression
von: Hill, Joshua, et al.
Veröffentlicht: (2025)
von: Hill, Joshua, et al.
Veröffentlicht: (2025)
Don't Forget to Connect! Improving RAG with Graph-based Reranking
von: Dong, Jialin, et al.
Veröffentlicht: (2024)
von: Dong, Jialin, et al.
Veröffentlicht: (2024)
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
von: Kim, Hyunseung, et al.
Veröffentlicht: (2024)
von: Kim, Hyunseung, et al.
Veröffentlicht: (2024)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
von: Zhang, Jiefu, et al.
Veröffentlicht: (2026)
von: Zhang, Jiefu, et al.
Veröffentlicht: (2026)
DESIRE: Dynamic Knowledge Consolidation for Rehearsal-Free Continual Learning
von: Guo, Haiyang, et al.
Veröffentlicht: (2024)
von: Guo, Haiyang, et al.
Veröffentlicht: (2024)
Don't Push the Button! Exploring Data Leakage Risks in Machine Learning and Transfer Learning
von: Apicella, Andrea, et al.
Veröffentlicht: (2024)
von: Apicella, Andrea, et al.
Veröffentlicht: (2024)
DSLR: Diversity Enhancement and Structure Learning for Rehearsal-based Graph Continual Learning
von: Choi, Seungyoon, et al.
Veröffentlicht: (2024)
von: Choi, Seungyoon, et al.
Veröffentlicht: (2024)
Task Scheduling & Forgetting in Multi-Task Reinforcement Learning
von: Speckmann, Marc, et al.
Veröffentlicht: (2025)
von: Speckmann, Marc, et al.
Veröffentlicht: (2025)
Pessimistic Value Iteration for Multi-Task Data Sharing in Offline Reinforcement Learning
von: Bai, Chenjia, et al.
Veröffentlicht: (2024)
von: Bai, Chenjia, et al.
Veröffentlicht: (2024)
Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target
von: Kim, Taesan, et al.
Veröffentlicht: (2026)
von: Kim, Taesan, et al.
Veröffentlicht: (2026)
Forget Forgetting: Continual Learning in a World of Abundant Memory
von: Cho, Dongkyu, et al.
Veröffentlicht: (2025)
von: Cho, Dongkyu, et al.
Veröffentlicht: (2025)
Keep Rehearsing and Refining: Lifelong Learning Vehicle Routing under Continually Drifting Tasks
von: Pei, Jiyuan, et al.
Veröffentlicht: (2026)
von: Pei, Jiyuan, et al.
Veröffentlicht: (2026)
CRAFT: Forgetting-Aware Intervention-Based Adaptation for Continual Learning
von: Hossen, Md Anwar, et al.
Veröffentlicht: (2026)
von: Hossen, Md Anwar, et al.
Veröffentlicht: (2026)
On the Implicit Adversariality of Catastrophic Forgetting in Deep Continual Learning
von: Peng, Ze, et al.
Veröffentlicht: (2025)
von: Peng, Ze, et al.
Veröffentlicht: (2025)
mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
von: Yang, Yongyi, et al.
Veröffentlicht: (2026)
von: Yang, Yongyi, et al.
Veröffentlicht: (2026)
Don't Shoot The Breeze: Topic Continuity Model Using Nonlinear Naive Bayes With Attention
von: Pi, Shu-Ting, et al.
Veröffentlicht: (2026)
von: Pi, Shu-Ting, et al.
Veröffentlicht: (2026)
PROL : Rehearsal Free Continual Learning in Streaming Data via Prompt Online Learning
von: Ma'sum, M. Anwar, et al.
Veröffentlicht: (2025)
von: Ma'sum, M. Anwar, et al.
Veröffentlicht: (2025)
Benchmarking is Broken -- Don't Let AI be its Own Judge
von: Cheng, Zerui, et al.
Veröffentlicht: (2025)
von: Cheng, Zerui, et al.
Veröffentlicht: (2025)
Sequencing to Mitigate Catastrophic Forgetting in Continual Learning
von: Moussa, Hesham G., et al.
Veröffentlicht: (2025)
von: Moussa, Hesham G., et al.
Veröffentlicht: (2025)
Accurate Forgetting for Heterogeneous Federated Continual Learning
von: Wuerkaixi, Abudukelimu, et al.
Veröffentlicht: (2025)
von: Wuerkaixi, Abudukelimu, et al.
Veröffentlicht: (2025)
Don't be lazy: CompleteP enables compute-efficient deep transformers
von: Dey, Nolan, et al.
Veröffentlicht: (2025)
von: Dey, Nolan, et al.
Veröffentlicht: (2025)
Don't Waste Your Time: Early Stopping Cross-Validation
von: Bergman, Edward, et al.
Veröffentlicht: (2024)
von: Bergman, Edward, et al.
Veröffentlicht: (2024)
DATA: Decomposed Attention-based Task Adaptation for Rehearsal-Free Continual Learning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
Multi-granularity Knowledge Transfer for Continual Reinforcement Learning
von: Pan, Chaofan, et al.
Veröffentlicht: (2024)
von: Pan, Chaofan, et al.
Veröffentlicht: (2024)
FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
von: Zhong, Shan, et al.
Veröffentlicht: (2025)
von: Zhong, Shan, et al.
Veröffentlicht: (2025)
Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding
von: Liu, Jiacheng, et al.
Veröffentlicht: (2023)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2023)
Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning
von: Elsayed, Mohamed, et al.
Veröffentlicht: (2024)
von: Elsayed, Mohamed, et al.
Veröffentlicht: (2024)
REP: Resource-Efficient Prompting for Rehearsal-Free Continual Learning
von: Jeon, Sungho, et al.
Veröffentlicht: (2024)
von: Jeon, Sungho, et al.
Veröffentlicht: (2024)
xAI-Drop: Don't Use What You Cannot Explain
von: De Luca, Vincenzo Marco, et al.
Veröffentlicht: (2024)
von: De Luca, Vincenzo Marco, et al.
Veröffentlicht: (2024)
Don't Lag, RAG: Training-Free Adversarial Detection Using RAG
von: Kazoom, Roie, et al.
Veröffentlicht: (2025)
von: Kazoom, Roie, et al.
Veröffentlicht: (2025)
Broad Critic Deep Actor Reinforcement Learning for Continuous Control
von: Thalagala, Shiron, et al.
Veröffentlicht: (2024)
von: Thalagala, Shiron, et al.
Veröffentlicht: (2024)
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
von: Shan, Zikang, et al.
Veröffentlicht: (2026)
von: Shan, Zikang, et al.
Veröffentlicht: (2026)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
von: Mayne, Harry, et al.
Veröffentlicht: (2025)
von: Mayne, Harry, et al.
Veröffentlicht: (2025)
Don't Play Favorites: Minority Guidance for Diffusion Models
von: Um, Soobin, et al.
Veröffentlicht: (2023)
von: Um, Soobin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Towards Interactive Reinforcement Learning with Intrinsic Feedback
von: Poole, Benjamin, et al.
Veröffentlicht: (2021) -
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026) -
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
von: Khoriaty, Matthew, et al.
Veröffentlicht: (2025) -
Don't Forget Imagination!
von: Vityaev, Evgenii E., et al.
Veröffentlicht: (2025) -
Continual Diffuser (CoD): Mastering Continual Offline Reinforcement Learning with Experience Rehearsal
von: Hu, Jifeng, et al.
Veröffentlicht: (2024)