Compute-Optimal Scaling for Value-Based Deep RL
Fuente:
arXiv
Guardado en:
| Autores principales: | Fu, Preston, Rybkin, Oleh, Zhou, Zhiyuan, Nauman, Michal, Abbeel, Pieter, Levine, Sergey, Kumar, Aviral |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Value-Based Deep RL Scales Predictably
por: Rybkin, Oleh, et al.
Publicado: (2025)
por: Rybkin, Oleh, et al.
Publicado: (2025)
floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
por: Agrawalla, Bhavya, et al.
Publicado: (2025)
por: Agrawalla, Bhavya, et al.
Publicado: (2025)
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
por: Nauman, Michal, et al.
Publicado: (2025)
por: Nauman, Michal, et al.
Publicado: (2025)
METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
por: Park, Seohong, et al.
Publicado: (2023)
por: Park, Seohong, et al.
Publicado: (2023)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
por: Setlur, Amrith, et al.
Publicado: (2025)
por: Setlur, Amrith, et al.
Publicado: (2025)
Is Value Learning Really the Main Bottleneck in Offline RL?
por: Park, Seohong, et al.
Publicado: (2024)
por: Park, Seohong, et al.
Publicado: (2024)
Reward-Conditioned Reinforcement Learning
por: Nauman, Michal, et al.
Publicado: (2026)
por: Nauman, Michal, et al.
Publicado: (2026)
Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos
por: Ye, Weirui, et al.
Publicado: (2025)
por: Ye, Weirui, et al.
Publicado: (2025)
A Stable Whitening Optimizer for Efficient Neural Network Training
por: Frans, Kevin, et al.
Publicado: (2025)
por: Frans, Kevin, et al.
Publicado: (2025)
What Really Matters in Matrix-Whitening Optimizers?
por: Frans, Kevin, et al.
Publicado: (2025)
por: Frans, Kevin, et al.
Publicado: (2025)
Cliqueformer: Model-Based Optimization with Structured Transformers
por: Kuba, Jakub Grudzien, et al.
Publicado: (2024)
por: Kuba, Jakub Grudzien, et al.
Publicado: (2024)
What Does Flow Matching Bring To TD Learning?
por: Agrawalla, Bhavya, et al.
Publicado: (2026)
por: Agrawalla, Bhavya, et al.
Publicado: (2026)
Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data
por: Zhou, Zhiyuan, et al.
Publicado: (2024)
por: Zhou, Zhiyuan, et al.
Publicado: (2024)
Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance
por: Nakamoto, Mitsuhiko, et al.
Publicado: (2024)
por: Nakamoto, Mitsuhiko, et al.
Publicado: (2024)
Diffusion Guidance Is a Controllable Policy Improvement Operator
por: Frans, Kevin, et al.
Publicado: (2025)
por: Frans, Kevin, et al.
Publicado: (2025)
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
por: Zhou, Yifei, et al.
Publicado: (2024)
por: Zhou, Yifei, et al.
Publicado: (2024)
When Does Non-Uniform Replay Matter in Reinforcement Learning?
por: Korniak, Michal, et al.
Publicado: (2026)
por: Korniak, Michal, et al.
Publicado: (2026)
Digi-Q: Learning Q-Value Functions for Training Device-Control Agents
por: Bai, Hao, et al.
Publicado: (2025)
por: Bai, Hao, et al.
Publicado: (2025)
Horizon Reduction Makes RL Scalable
por: Park, Seohong, et al.
Publicado: (2025)
por: Park, Seohong, et al.
Publicado: (2025)
Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings
por: Frans, Kevin, et al.
Publicado: (2024)
por: Frans, Kevin, et al.
Publicado: (2024)
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
por: Bai, Hao, et al.
Publicado: (2024)
por: Bai, Hao, et al.
Publicado: (2024)
One Step Diffusion via Shortcut Models
por: Frans, Kevin, et al.
Publicado: (2024)
por: Frans, Kevin, et al.
Publicado: (2024)
Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
por: Farebrother, Jesse, et al.
Publicado: (2024)
por: Farebrother, Jesse, et al.
Publicado: (2024)
Functional Graphical Models: Structure Enables Offline Data-Driven Optimization
por: Kuba, Jakub Grudzien, et al.
Publicado: (2024)
por: Kuba, Jakub Grudzien, et al.
Publicado: (2024)
IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL
por: Cheng, Zhoujun, et al.
Publicado: (2026)
por: Cheng, Zhoujun, et al.
Publicado: (2026)
Prioritized Generative Replay
por: Wang, Renhao, et al.
Publicado: (2024)
por: Wang, Renhao, et al.
Publicado: (2024)
SOMBRL: Scalable and Optimistic Model-Based RL
por: Sukhija, Bhavya, et al.
Publicado: (2025)
por: Sukhija, Bhavya, et al.
Publicado: (2025)
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
por: Snell, Charlie, et al.
Publicado: (2024)
por: Snell, Charlie, et al.
Publicado: (2024)
FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
por: Seo, Younggyo, et al.
Publicado: (2025)
por: Seo, Younggyo, et al.
Publicado: (2025)
Transitive RL: Value Learning via Divide and Conquer
por: Park, Seohong, et al.
Publicado: (2025)
por: Park, Seohong, et al.
Publicado: (2025)
ViVa: Video-Trained Value Functions for Guiding Online RL from Diverse Data
por: Dashora, Nitish, et al.
Publicado: (2025)
por: Dashora, Nitish, et al.
Publicado: (2025)
D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
por: Rafailov, Rafael, et al.
Publicado: (2024)
por: Rafailov, Rafael, et al.
Publicado: (2024)
Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning
por: Nakamoto, Mitsuhiko, et al.
Publicado: (2023)
por: Nakamoto, Mitsuhiko, et al.
Publicado: (2023)
Vision-Language Models Provide Promptable Representations for Reinforcement Learning
por: Chen, William, et al.
Publicado: (2024)
por: Chen, William, et al.
Publicado: (2024)
Privileged Sensing Scaffolds Reinforcement Learning
por: Hu, Edward S., et al.
Publicado: (2024)
por: Hu, Edward S., et al.
Publicado: (2024)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
por: Kim, Dongyoung, et al.
Publicado: (2023)
por: Kim, Dongyoung, et al.
Publicado: (2023)
Scalable Offline Model-Based RL with Action Chunks
por: Park, Kwanyoung, et al.
Publicado: (2025)
por: Park, Kwanyoung, et al.
Publicado: (2025)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
por: Ankile, Lars, et al.
Publicado: (2025)
por: Ankile, Lars, et al.
Publicado: (2025)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
por: Wu, Mian, et al.
Publicado: (2025)
por: Wu, Mian, et al.
Publicado: (2025)
Unfamiliar Finetuning Examples Control How Language Models Hallucinate
por: Kang, Katie, et al.
Publicado: (2024)
por: Kang, Katie, et al.
Publicado: (2024)
Ejemplares similares
-
Value-Based Deep RL Scales Predictably
por: Rybkin, Oleh, et al.
Publicado: (2025) -
floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
por: Agrawalla, Bhavya, et al.
Publicado: (2025) -
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
por: Nauman, Michal, et al.
Publicado: (2025) -
METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
por: Park, Seohong, et al.
Publicado: (2023) -
Scaling Test-Time Compute Without Verification or RL is Suboptimal
por: Setlur, Amrith, et al.
Publicado: (2025)