Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
Fuente:
arXiv
Saved in:
| Main Authors: | Nauman, Michal, Cygan, Marek, Sferrazza, Carmelo, Kumar, Aviral, Abbeel, Pieter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reward-Conditioned Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2026)
by: Nauman, Michal, et al.
Published: (2026)
Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
On the Theory of Risk-Aware Agents: Bridging Actor-Critic and Economics
by: Nauman, Michal, et al.
Published: (2023)
by: Nauman, Michal, et al.
Published: (2023)
FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
by: Seo, Younggyo, et al.
Published: (2025)
by: Seo, Younggyo, et al.
Published: (2025)
Value-Based Deep RL Scales Predictably
by: Rybkin, Oleh, et al.
Published: (2025)
by: Rybkin, Oleh, et al.
Published: (2025)
Compute-Optimal Scaling for Value-Based Deep RL
by: Fu, Preston, et al.
Published: (2025)
by: Fu, Preston, et al.
Published: (2025)
A Case for Validation Buffer in Pessimistic Actor-Critic
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
by: Sukhija, Bhavya, et al.
Published: (2024)
by: Sukhija, Bhavya, et al.
Published: (2024)
floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
by: Agrawalla, Bhavya, et al.
Published: (2025)
by: Agrawalla, Bhavya, et al.
Published: (2025)
SOMBRL: Scalable and Optimistic Model-Based RL
by: Sukhija, Bhavya, et al.
Published: (2025)
by: Sukhija, Bhavya, et al.
Published: (2025)
HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation
by: Sferrazza, Carmelo, et al.
Published: (2024)
by: Sferrazza, Carmelo, et al.
Published: (2024)
Body Transformer: Leveraging Robot Embodiment for Policy Learning
by: Sferrazza, Carmelo, et al.
Published: (2024)
by: Sferrazza, Carmelo, et al.
Published: (2024)
What Does Flow Matching Bring To TD Learning?
by: Agrawalla, Bhavya, et al.
Published: (2026)
by: Agrawalla, Bhavya, et al.
Published: (2026)
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
Learning Sim-to-Real Humanoid Locomotion in 15 Minutes
by: Seo, Younggyo, et al.
Published: (2025)
by: Seo, Younggyo, et al.
Published: (2025)
When Does Non-Uniform Replay Matter in Reinforcement Learning?
by: Korniak, Michal, et al.
Published: (2026)
by: Korniak, Michal, et al.
Published: (2026)
Debate2Create: Robot Co-design via Multi-Agent LLM Debate
by: Qiu, Kevin, et al.
Published: (2025)
by: Qiu, Kevin, et al.
Published: (2025)
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024)
by: Seo, Younggyo, et al.
Published: (2024)
A Stable Whitening Optimizer for Efficient Neural Network Training
by: Frans, Kevin, et al.
Published: (2025)
by: Frans, Kevin, et al.
Published: (2025)
OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
by: Yang, Lujie, et al.
Published: (2025)
by: Yang, Lujie, et al.
Published: (2025)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
by: Kim, Dongyoung, et al.
Published: (2023)
by: Kim, Dongyoung, et al.
Published: (2023)
Relative Entropy Pathwise Policy Optimization
by: Voelcker, Claas, et al.
Published: (2025)
by: Voelcker, Claas, et al.
Published: (2025)
Digi-Q: Learning Q-Value Functions for Training Device-Control Agents
by: Bai, Hao, et al.
Published: (2025)
by: Bai, Hao, et al.
Published: (2025)
Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings
by: Frans, Kevin, et al.
Published: (2024)
by: Frans, Kevin, et al.
Published: (2024)
What Really Matters in Matrix-Whitening Optimizers?
by: Frans, Kevin, et al.
Published: (2025)
by: Frans, Kevin, et al.
Published: (2025)
SEMDICE: Off-policy State Entropy Maximization via Stationary Distribution Correction Estimation
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
What Matters in Hierarchical Search for Combinatorial Reasoning Problems?
by: Zawalski, Michał, et al.
Published: (2024)
by: Zawalski, Michał, et al.
Published: (2024)
Offline Imitation Learning Through Graph Search and Retrieval
by: Yin, Zhao-Heng, et al.
Published: (2024)
by: Yin, Zhao-Heng, et al.
Published: (2024)
Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching
by: Wu, Zhen, et al.
Published: (2026)
by: Wu, Zhen, et al.
Published: (2026)
Functional Graphical Models: Structure Enables Offline Data-Driven Optimization
by: Kuba, Jakub Grudzien, et al.
Published: (2024)
by: Kuba, Jakub Grudzien, et al.
Published: (2024)
Vid2Sid: Videos Can Help Close the Sim2Real Gap
by: Qiu, Kevin, et al.
Published: (2026)
by: Qiu, Kevin, et al.
Published: (2026)
ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning
by: As, Yarden, et al.
Published: (2024)
by: As, Yarden, et al.
Published: (2024)
Is Value Learning Really the Main Bottleneck in Offline RL?
by: Park, Seohong, et al.
Published: (2024)
by: Park, Seohong, et al.
Published: (2024)
Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance
by: Nakamoto, Mitsuhiko, et al.
Published: (2024)
by: Nakamoto, Mitsuhiko, et al.
Published: (2024)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
by: Lee, Vint, et al.
Published: (2023)
by: Lee, Vint, et al.
Published: (2023)
Diffusion Guidance Is a Controllable Policy Improvement Operator
by: Frans, Kevin, et al.
Published: (2025)
by: Frans, Kevin, et al.
Published: (2025)
World Model on Million-Length Video And Language With Blockwise RingAttention
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
Cliqueformer: Model-Based Optimization with Structured Transformers
by: Kuba, Jakub Grudzien, et al.
Published: (2024)
by: Kuba, Jakub Grudzien, et al.
Published: (2024)
Projected Compression: Trainable Projection for Efficient Transformer Compression
by: Stefaniak, Maciej, et al.
Published: (2025)
by: Stefaniak, Maciej, et al.
Published: (2025)
Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos
by: Ye, Weirui, et al.
Published: (2025)
by: Ye, Weirui, et al.
Published: (2025)
Similar Items
-
Reward-Conditioned Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2026) -
Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
by: Nauman, Michal, et al.
Published: (2024) -
On the Theory of Risk-Aware Agents: Bridging Actor-Critic and Economics
by: Nauman, Michal, et al.
Published: (2023) -
FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
by: Seo, Younggyo, et al.
Published: (2025) -
Value-Based Deep RL Scales Predictably
by: Rybkin, Oleh, et al.
Published: (2025)