Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
Fuente:
arXiv
Salvato in:
| Autori principali: | Clark, Tyler, Evers, Christine, Hare, Jonathon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond The Rainbow: High Performance Deep Reinforcement Learning on a Desktop PC
di: Clark, Tyler, et al.
Pubblicazione: (2024)
di: Clark, Tyler, et al.
Pubblicazione: (2024)
Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
di: Dang, Quy-Anh, et al.
Pubblicazione: (2025)
di: Dang, Quy-Anh, et al.
Pubblicazione: (2025)
When More Data Doesn't Help: Limits of Adaptation in Multitask Learning
di: Hanneke, Steve, et al.
Pubblicazione: (2026)
di: Hanneke, Steve, et al.
Pubblicazione: (2026)
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
di: Falahati, Ali, et al.
Pubblicazione: (2026)
di: Falahati, Ali, et al.
Pubblicazione: (2026)
Semantics at an Angle: When Cosine Similarity Works Until It Doesn't
di: You, Kisung
Pubblicazione: (2025)
di: You, Kisung
Pubblicazione: (2025)
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
di: Berlot-Attwell, Ian, et al.
Pubblicazione: (2024)
di: Berlot-Attwell, Ian, et al.
Pubblicazione: (2024)
Rethinking Deep Thinking: Stable Learning of Algorithms using Lipschitz Constraints
di: Bear, Jay, et al.
Pubblicazione: (2024)
di: Bear, Jay, et al.
Pubblicazione: (2024)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
di: Cooper, A. Feder, et al.
Pubblicazione: (2024)
di: Cooper, A. Feder, et al.
Pubblicazione: (2024)
A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't)
di: Nayak, Nihal V., et al.
Pubblicazione: (2026)
di: Nayak, Nihal V., et al.
Pubblicazione: (2026)
Infinite Width Models That Work: Why Feature Learning Doesn't Matter as Much as You Think
di: Sernau, Luke
Pubblicazione: (2024)
di: Sernau, Luke
Pubblicazione: (2024)
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
di: He, Di, et al.
Pubblicazione: (2026)
di: He, Di, et al.
Pubblicazione: (2026)
Explorations of the Softmax Space: Knowing When the Neural Network Doesn't Know
di: Sikar, Daniel, et al.
Pubblicazione: (2025)
di: Sikar, Daniel, et al.
Pubblicazione: (2025)
Time-Warping Recurrent Neural Networks for Transfer Learning
di: Hirschi, Jonathon
Pubblicazione: (2026)
di: Hirschi, Jonathon
Pubblicazione: (2026)
When Structure Doesn't Help: LLMs Do Not Read Text-Attributed Graphs as Effectively as We Expected
di: Xu, Haotian, et al.
Pubblicazione: (2025)
di: Xu, Haotian, et al.
Pubblicazione: (2025)
Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies
di: Liu, Ming
Pubblicazione: (2026)
di: Liu, Ming
Pubblicazione: (2026)
On the Reuse Bias in Off-Policy Reinforcement Learning
di: Ying, Chengyang, et al.
Pubblicazione: (2022)
di: Ying, Chengyang, et al.
Pubblicazione: (2022)
MedCalc-Bench Doesn't Measure What You Think: A Benchmark Audit and the Case for Open-Book Evaluation
di: Krohn-Grimberghe, Artus
Pubblicazione: (2026)
di: Krohn-Grimberghe, Artus
Pubblicazione: (2026)
Doesn't Everyone Have Rights to a Learner's Permit?
di: Gehrig, Jody, et al.
Pubblicazione: (2009)
di: Gehrig, Jody, et al.
Pubblicazione: (2009)
Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis
di: Gong, Shuzhi, et al.
Pubblicazione: (2026)
di: Gong, Shuzhi, et al.
Pubblicazione: (2026)
Off-Policy Primal-Dual Safe Reinforcement Learning
di: Wu, Zifan, et al.
Pubblicazione: (2024)
di: Wu, Zifan, et al.
Pubblicazione: (2024)
Multi-Agent Deep Reinforcement Learning Under Constrained Communications
di: Shaik, Shahil, et al.
Pubblicazione: (2026)
di: Shaik, Shahil, et al.
Pubblicazione: (2026)
Action-Graph Policies: Learning Action Co-dependencies in Multi-Agent Reinforcement Learning
di: Gupta, Nikunj, et al.
Pubblicazione: (2026)
di: Gupta, Nikunj, et al.
Pubblicazione: (2026)
RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
di: Hisaki, Yukinari, et al.
Pubblicazione: (2024)
di: Hisaki, Yukinari, et al.
Pubblicazione: (2024)
Revisiting Intermediate-Layer Matching in Knowledge Distillation: Layer-Selection Strategy Doesn't Matter (Much)
di: Yu, Zony, et al.
Pubblicazione: (2025)
di: Yu, Zony, et al.
Pubblicazione: (2025)
Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement Learning
di: Daley, Brett, et al.
Pubblicazione: (2023)
di: Daley, Brett, et al.
Pubblicazione: (2023)
Deep Meta Coordination Graphs for Multi-agent Reinforcement Learning
di: Gupta, Nikunj, et al.
Pubblicazione: (2025)
di: Gupta, Nikunj, et al.
Pubblicazione: (2025)
Off Policy Lyapunov Stability in Reinforcement Learning
di: Gill, Sarvan, et al.
Pubblicazione: (2025)
di: Gill, Sarvan, et al.
Pubblicazione: (2025)
Efficient Recurrent Off-Policy RL Requires a Context-Encoder-Specific Learning Rate
di: Luo, Fan-Ming, et al.
Pubblicazione: (2024)
di: Luo, Fan-Ming, et al.
Pubblicazione: (2024)
Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
di: Xiao, Teng, et al.
Pubblicazione: (2024)
di: Xiao, Teng, et al.
Pubblicazione: (2024)
CUER: Corrected Uniform Experience Replay for Off-Policy Continuous Deep Reinforcement Learning Algorithms
di: Yenicesu, Arda Sarp, et al.
Pubblicazione: (2024)
di: Yenicesu, Arda Sarp, et al.
Pubblicazione: (2024)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
di: Li, Guopeng, et al.
Pubblicazione: (2026)
di: Li, Guopeng, et al.
Pubblicazione: (2026)
Transfer Learning for Dead Fuel Moisture Prediction Using Time-Warping Recurrent Neural Networks
di: Hirschi, Jonathon, et al.
Pubblicazione: (2026)
di: Hirschi, Jonathon, et al.
Pubblicazione: (2026)
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
di: Palenicek, Daniel, et al.
Pubblicazione: (2025)
di: Palenicek, Daniel, et al.
Pubblicazione: (2025)
Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning
di: Zhuang, Yuan, et al.
Pubblicazione: (2026)
di: Zhuang, Yuan, et al.
Pubblicazione: (2026)
TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement Learning
di: Li, Ge, et al.
Pubblicazione: (2024)
di: Li, Ge, et al.
Pubblicazione: (2024)
Consciousness Doesn't Do That
di: Matthias Michel
Pubblicazione: (2026)
di: Matthias Michel
Pubblicazione: (2026)
"Something Comes through or It Doesn't": Intensive Reading in Post-Qualitative Inquiry
di: Maggie MacLure
Pubblicazione: (2024)
di: Maggie MacLure
Pubblicazione: (2024)
Neural Co-state Policies: Structuring Hidden States in Recurrent Reinforcement Learning
di: Leeftink, David, et al.
Pubblicazione: (2026)
di: Leeftink, David, et al.
Pubblicazione: (2026)
Succeed or Learn Slowly: Sample Efficient Off-Policy Reinforcement Learning for Mobile App Control
di: Papoudakis, Georgios, et al.
Pubblicazione: (2025)
di: Papoudakis, Georgios, et al.
Pubblicazione: (2025)
Koopman Regularized Deep Speech Disentanglement for Speaker Verification
di: Chazaridis, Nikos, et al.
Pubblicazione: (2026)
di: Chazaridis, Nikos, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Beyond The Rainbow: High Performance Deep Reinforcement Learning on a Desktop PC
di: Clark, Tyler, et al.
Pubblicazione: (2024) -
Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
di: Dang, Quy-Anh, et al.
Pubblicazione: (2025) -
When More Data Doesn't Help: Limits of Adaptation in Multitask Learning
di: Hanneke, Steve, et al.
Pubblicazione: (2026) -
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
di: Falahati, Ali, et al.
Pubblicazione: (2026) -
Semantics at an Angle: When Cosine Similarity Works Until It Doesn't
di: You, Kisung
Pubblicazione: (2025)