Massively Scaling Explicit Policy-conditioned Value Functions
Fuente:
arXiv
Saved in:
| Main Authors: | Bohlinger, Nico, Peters, Jan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shape Your Body: Value Gradients for Multi-Embodiment Robot Design
by: Bohlinger, Nico, et al.
Published: (2026)
by: Bohlinger, Nico, et al.
Published: (2026)
Multi-Embodiment Locomotion at Scale with extreme Embodiment Randomization
by: Bohlinger, Nico, et al.
Published: (2025)
by: Bohlinger, Nico, et al.
Published: (2025)
Towards Embodiment Scaling Laws in Robot Locomotion
by: Ai, Bo, et al.
Published: (2025)
by: Ai, Bo, et al.
Published: (2025)
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
Model-Based Epistemic Variance of Values for Risk-Aware Policy Optimization
by: Luis, Carlos E., et al.
Published: (2023)
by: Luis, Carlos E., et al.
Published: (2023)
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
by: Diwan, Anish, et al.
Published: (2026)
by: Diwan, Anish, et al.
Published: (2026)
Scaling CrossQ with Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
Value-Distributional Model-Based Reinforcement Learning
by: Luis, Carlos E., et al.
Published: (2023)
by: Luis, Carlos E., et al.
Published: (2023)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
by: Reddi, Aryaman, et al.
Published: (2025)
by: Reddi, Aryaman, et al.
Published: (2025)
XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
by: Huang, Chenghua, et al.
Published: (2025)
by: Huang, Chenghua, et al.
Published: (2025)
Discrete Variational Autoencoding via Policy Search
by: Drolet, Michael, et al.
Published: (2025)
by: Drolet, Michael, et al.
Published: (2025)
Gait in Eight: Efficient On-Robot Learning for Omnidirectional Quadruped Locomotion
by: Bohlinger, Nico, et al.
Published: (2025)
by: Bohlinger, Nico, et al.
Published: (2025)
Per-Domain Generalizing Policies: On Learning Efficient and Robust Q-Value Functions (Extended Version with Technical Appendix)
by: Müller, Nicola J., et al.
Published: (2026)
by: Müller, Nicola J., et al.
Published: (2026)
One Policy to Run Them All: an End-to-end Learning Approach to Multi-Embodiment Locomotion
by: Bohlinger, Nico, et al.
Published: (2024)
by: Bohlinger, Nico, et al.
Published: (2024)
Universal Value-Function Uncertainties
by: Zanger, Moritz A., et al.
Published: (2025)
by: Zanger, Moritz A., et al.
Published: (2025)
Value-Free Policy Optimization via Reward Partitioning
by: Faye, Bilal, et al.
Published: (2025)
by: Faye, Bilal, et al.
Published: (2025)
Quasimetric Value Functions with Dense Rewards
by: Valieva, Khadichabonu, et al.
Published: (2024)
by: Valieva, Khadichabonu, et al.
Published: (2024)
C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions
by: Kubo, Kenji, et al.
Published: (2026)
by: Kubo, Kenji, et al.
Published: (2026)
Scaling Policy Gradient Quality-Diversity with Massive Parallelization via Behavioral Variations
by: Mitsides, Konstantinos, et al.
Published: (2025)
by: Mitsides, Konstantinos, et al.
Published: (2025)
Peer Learning: Learning Complex Policies in Groups from Scratch via Action Recommendations
by: Derstroff, Cedric, et al.
Published: (2023)
by: Derstroff, Cedric, et al.
Published: (2023)
Inferring Transition Dynamics from Value Functions
by: Adamczyk, Jacob
Published: (2025)
by: Adamczyk, Jacob
Published: (2025)
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
Residual Q-Learning: Offline and Online Policy Customization without Value
by: Li, Chenran, et al.
Published: (2023)
by: Li, Chenran, et al.
Published: (2023)
Predicting Long Term Sequential Policy Value Using Softer Surrogates
by: Nam, Hyunji, et al.
Published: (2024)
by: Nam, Hyunji, et al.
Published: (2024)
Functional Acceleration for Policy Mirror Descent
by: Chelu, Veronica, et al.
Published: (2024)
by: Chelu, Veronica, et al.
Published: (2024)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
by: Zhu, Dingwei, et al.
Published: (2025)
by: Zhu, Dingwei, et al.
Published: (2025)
Improving Deep Reinforcement Learning by Reducing the Chain Effect of Value and Policy Churn
by: Tang, Hongyao, et al.
Published: (2024)
by: Tang, Hongyao, et al.
Published: (2024)
M$^2$OE$^2$-GL: A Family of Probabilistic Load Forecasters That Scales to Massive Customers
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
LLM4XCE: Large Language Models for Extremely Large-Scale Massive MIMO Channel Estimation
by: Li, Renbin, et al.
Published: (2025)
by: Li, Renbin, et al.
Published: (2025)
Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures
by: Wang, Haohui, et al.
Published: (2025)
by: Wang, Haohui, et al.
Published: (2025)
VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning
by: Chen, Xuyang, et al.
Published: (2025)
by: Chen, Xuyang, et al.
Published: (2025)
Adaptive Exploration for Data-Efficient General Value Function Evaluations
by: Jain, Arushi, et al.
Published: (2024)
by: Jain, Arushi, et al.
Published: (2024)
Stable Offline Value Function Learning with Bisimulation-based Representations
by: Pavse, Brahma S., et al.
Published: (2024)
by: Pavse, Brahma S., et al.
Published: (2024)
Tensor Low-rank Approximation of Finite-horizon Value Functions
by: Rozada, Sergio, et al.
Published: (2024)
by: Rozada, Sergio, et al.
Published: (2024)
Neuro-Symbolic Imitation Learning: Discovering Symbolic Abstractions for Skill Learning
by: Keller, Leon, et al.
Published: (2025)
by: Keller, Leon, et al.
Published: (2025)
Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement Learning
by: Tiofack, Franki Nguimatsia, et al.
Published: (2025)
by: Tiofack, Franki Nguimatsia, et al.
Published: (2025)
Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning
by: Ma, Yunchang, et al.
Published: (2025)
by: Ma, Yunchang, et al.
Published: (2025)
Backward Learning for Goal-Conditioned Policies
by: Höftmann, Marc, et al.
Published: (2023)
by: Höftmann, Marc, et al.
Published: (2023)
Smooth Gate Functions for Soft Advantage Policy Optimization
by: Denisov, Egor, et al.
Published: (2026)
by: Denisov, Egor, et al.
Published: (2026)
Similar Items
-
Shape Your Body: Value Gradients for Multi-Embodiment Robot Design
by: Bohlinger, Nico, et al.
Published: (2026) -
Multi-Embodiment Locomotion at Scale with extreme Embodiment Randomization
by: Bohlinger, Nico, et al.
Published: (2025) -
Towards Embodiment Scaling Laws in Robot Locomotion
by: Ai, Bo, et al.
Published: (2025) -
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025) -
Model-Based Epistemic Variance of Values for Risk-Aware Policy Optimization
by: Luis, Carlos E., et al.
Published: (2023)