Towards General-Purpose Model-Free Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fujimoto, Scott, D'Oro, Pierluca, Zhang, Amy, Tian, Yuandong, Rabbat, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hierarchical Behaviour Spaces
von: Matthews, Michael Tryfan, et al.
Veröffentlicht: (2026)
von: Matthews, Michael Tryfan, et al.
Veröffentlicht: (2026)
Do Transformer World Models Give Better Policy Gradients?
von: Ma, Michel, et al.
Veröffentlicht: (2024)
von: Ma, Michel, et al.
Veröffentlicht: (2024)
ADEPTS: A Capability Framework for Human-Centered Agent Design
von: D'Oro, Pierluca, et al.
Veröffentlicht: (2025)
von: D'Oro, Pierluca, et al.
Veröffentlicht: (2025)
Scalable Option Learning in High-Throughput Environments
von: Henaff, Mikael, et al.
Veröffentlicht: (2025)
von: Henaff, Mikael, et al.
Veröffentlicht: (2025)
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
Maxwell's Demon at Work: Efficient Pruning by Leveraging Saturation of Neurons
von: Dufort-Labbé, Simon, et al.
Veröffentlicht: (2024)
von: Dufort-Labbé, Simon, et al.
Veröffentlicht: (2024)
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces
von: Su, DiJia, et al.
Veröffentlicht: (2024)
von: Su, DiJia, et al.
Veröffentlicht: (2024)
Mol-MoE: Training Preference-Guided Routers for Molecule Generation
von: Calanzone, Diego, et al.
Veröffentlicht: (2025)
von: Calanzone, Diego, et al.
Veröffentlicht: (2025)
Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking
von: Tian, Yuandong
Veröffentlicht: (2025)
von: Tian, Yuandong
Veröffentlicht: (2025)
Controlling Multimodal LLMs via Reward-guided Decoding
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
MaestroMotif: Skill Design from Artificial Intelligence Feedback
von: Klissarov, Martin, et al.
Veröffentlicht: (2024)
von: Klissarov, Martin, et al.
Veröffentlicht: (2024)
Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity
von: Arnob, Samin Yeasar, et al.
Veröffentlicht: (2025)
von: Arnob, Samin Yeasar, et al.
Veröffentlicht: (2025)
The Curse of Diversity in Ensemble-Based Exploration
von: Lin, Zhixuan, et al.
Veröffentlicht: (2024)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2024)
DigiData: Training and Evaluating General-Purpose Mobile Control Agents
von: Sun, Yuxuan, et al.
Veröffentlicht: (2025)
von: Sun, Yuxuan, et al.
Veröffentlicht: (2025)
A Principled Loss Function for Direct Language Model Alignment
von: Tan, Yuandong
Veröffentlicht: (2025)
von: Tan, Yuandong
Veröffentlicht: (2025)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
von: Xu, Haoran, et al.
Veröffentlicht: (2025)
von: Xu, Haoran, et al.
Veröffentlicht: (2025)
Multi-objective Optimization by Learning Space Partitions
von: Zhao, Yiyang, et al.
Veröffentlicht: (2021)
von: Zhao, Yiyang, et al.
Veröffentlicht: (2021)
Towards General Purpose Robots at Scale: Lifelong Learning and Learning to Use Memory
von: Yue, William
Veröffentlicht: (2024)
von: Yue, William
Veröffentlicht: (2024)
SMORE: Score Models for Offline Goal-Conditioned Reinforcement Learning
von: Sikchi, Harshit, et al.
Veröffentlicht: (2023)
von: Sikchi, Harshit, et al.
Veröffentlicht: (2023)
Benchmarking General-Purpose In-Context Learning
von: Wang, Fan, et al.
Veröffentlicht: (2024)
von: Wang, Fan, et al.
Veröffentlicht: (2024)
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
von: Sikchi, Harshit, et al.
Veröffentlicht: (2023)
von: Sikchi, Harshit, et al.
Veröffentlicht: (2023)
Towards Robust Offline Reinforcement Learning under Diverse Data Corruption
von: Yang, Rui, et al.
Veröffentlicht: (2023)
von: Yang, Rui, et al.
Veröffentlicht: (2023)
Offline-to-Online Reinforcement Learning with Classifier-Free Diffusion Generation
von: Huang, Xiao, et al.
Veröffentlicht: (2025)
von: Huang, Xiao, et al.
Veröffentlicht: (2025)
Positive Unlabeled Contrastive Learning
von: Acharya, Anish, et al.
Veröffentlicht: (2022)
von: Acharya, Anish, et al.
Veröffentlicht: (2022)
Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning
von: Chuck, Caleb, et al.
Veröffentlicht: (2025)
von: Chuck, Caleb, et al.
Veröffentlicht: (2025)
Composing Global Solutions to Reasoning Tasks via Algebraic Objects in Neural Nets
von: Tian, Yuandong
Veröffentlicht: (2024)
von: Tian, Yuandong
Veröffentlicht: (2024)
On Zero-Shot Reinforcement Learning
von: Jeen, Scott
Veröffentlicht: (2025)
von: Jeen, Scott
Veröffentlicht: (2025)
Zero-Shot Reinforcement Learning via Function Encoders
von: Ingebrand, Tyler, et al.
Veröffentlicht: (2024)
von: Ingebrand, Tyler, et al.
Veröffentlicht: (2024)
Towards a Scalable Reference-Free Evaluation of Generative Models
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
Evolutionary Discovery of Reinforcement Learning Algorithms via Large Language Models
von: Sygkounas, Alkis, et al.
Veröffentlicht: (2026)
von: Sygkounas, Alkis, et al.
Veröffentlicht: (2026)
Structure in Deep Reinforcement Learning: A Survey and Open Problems
von: Mohan, Aditya, et al.
Veröffentlicht: (2023)
von: Mohan, Aditya, et al.
Veröffentlicht: (2023)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
von: Bush, Thomas, et al.
Veröffentlicht: (2025)
von: Bush, Thomas, et al.
Veröffentlicht: (2025)
Towards Interpretable Deep Reinforcement Learning Models via Inverse Reinforcement Learning
von: Xie, Sean, et al.
Veröffentlicht: (2022)
von: Xie, Sean, et al.
Veröffentlicht: (2022)
FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model
von: Gao, Chongkai, et al.
Veröffentlicht: (2024)
von: Gao, Chongkai, et al.
Veröffentlicht: (2024)
Reinforcement Learning via Value Gradient Flow
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
von: Rashidinejad, Paria, et al.
Veröffentlicht: (2024)
von: Rashidinejad, Paria, et al.
Veröffentlicht: (2024)
Towards Monotonic Improvement in In-Context Reinforcement Learning
von: Zhang, Wenhao, et al.
Veröffentlicht: (2025)
von: Zhang, Wenhao, et al.
Veröffentlicht: (2025)
Model-Free Robust Reinforcement Learning with Sample Complexity Analysis
von: Wang, Yudan, et al.
Veröffentlicht: (2024)
von: Wang, Yudan, et al.
Veröffentlicht: (2024)
Label-Free Reinforcement Learning via Cross-Model Entropy
von: Gorbett, Matt, et al.
Veröffentlicht: (2026)
von: Gorbett, Matt, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Hierarchical Behaviour Spaces
von: Matthews, Michael Tryfan, et al.
Veröffentlicht: (2026) -
Do Transformer World Models Give Better Policy Gradients?
von: Ma, Michel, et al.
Veröffentlicht: (2024) -
ADEPTS: A Capability Framework for Human-Centered Agent Design
von: D'Oro, Pierluca, et al.
Veröffentlicht: (2025) -
Scalable Option Learning in High-Throughput Environments
von: Henaff, Mikael, et al.
Veröffentlicht: (2025) -
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
von: Ding, Zihan, et al.
Veröffentlicht: (2024)