When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Dodeja, Lakshita, Biza, Ondrej, Vats, Shivam, Hart, Stephen, Tellex, Stefanie, Walters, Robin, Schmeckpeper, Karl, Weng, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating Residual Reinforcement Learning with Uncertainty Estimation
by: Dodeja, Lakshita, et al.
Published: (2025)
by: Dodeja, Lakshita, et al.
Published: (2025)
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning
by: Schroeder, Philip, et al.
Published: (2026)
by: Schroeder, Philip, et al.
Published: (2026)
You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector
by: Patil, Omkar, et al.
Published: (2026)
by: Patil, Omkar, et al.
Published: (2026)
Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies
by: Huang, Haojie, et al.
Published: (2024)
by: Huang, Haojie, et al.
Published: (2024)
On-Robot Reinforcement Learning with Goal-Contrastive Rewards
by: Biza, Ondrej, et al.
Published: (2024)
by: Biza, Ondrej, et al.
Published: (2024)
ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
by: Schroeder, Philip, et al.
Published: (2025)
by: Schroeder, Philip, et al.
Published: (2025)
One-Shot Cross-Geometry Skill Transfer through Part Decomposition
by: Thompson, Skye, et al.
Published: (2026)
by: Thompson, Skye, et al.
Published: (2026)
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs
by: Khairi, Ammar, et al.
Published: (2025)
by: Khairi, Ammar, et al.
Published: (2025)
Stable-BC: Controlling Covariate Shift with Stable Behavior Cloning
by: Mehta, Shaunak A., et al.
Published: (2024)
by: Mehta, Shaunak A., et al.
Published: (2024)
SelfBC: Self Behavior Cloning for Offline Reinforcement Learning
by: Liu, Shirong, et al.
Published: (2024)
by: Liu, Shirong, et al.
Published: (2024)
Research on Teaching and Learning Mathematics at the Tertiary Level State-of-the-art and Looking Ahead
by: Irene Biza
by: Irene Biza
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
by: Wagenmaker, Andrew, et al.
Published: (2025)
by: Wagenmaker, Andrew, et al.
Published: (2025)
Properties of $Q^{5}q$ dibaryons
by: Weng, Xin-Zhen
Published: (2024)
by: Weng, Xin-Zhen
Published: (2024)
Q&A: Michael Honey
by: Helicher, Karl
Published: (2007)
by: Helicher, Karl
Published: (2007)
Kibble-Zurek scalings and coarsening laws in slowly quenched classical Ising chains
by: Jindal, Lakshita, et al.
Published: (2024)
by: Jindal, Lakshita, et al.
Published: (2024)
Kink-kink correlations in nonlinear quenches across a quantum critical point
by: Jindal, Lakshita, et al.
Published: (2026)
by: Jindal, Lakshita, et al.
Published: (2026)
Scaling regimes in slow quenches within a gapped phase
by: Jindal, Lakshita, et al.
Published: (2025)
by: Jindal, Lakshita, et al.
Published: (2025)
Advanced Chest X-Ray Analysis via Transformer-Based Image Descriptors and Cross-Model Attention Mechanism
by: Agarwal, Lakshita, et al.
Published: (2025)
by: Agarwal, Lakshita, et al.
Published: (2025)
Towards Explainable AI: Multi-Modal Transformer for Video-based Image Description Generation
by: Agarwal, Lakshita, et al.
Published: (2025)
by: Agarwal, Lakshita, et al.
Published: (2025)
Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism
by: Agarwal, Lakshita, et al.
Published: (2025)
by: Agarwal, Lakshita, et al.
Published: (2025)
Isomorphisms and automorphisms of multiprojective bundles and symmetric powers of projective bundles
by: Bansal, Ashima, et al.
Published: (2025)
by: Bansal, Ashima, et al.
Published: (2025)
Automorphisms of punctual Hilbert schemes and symmetric powers of varieties
by: Bansal, Ashima, et al.
Published: (2025)
by: Bansal, Ashima, et al.
Published: (2025)
Singularity of cubic hypersurfaces and hyperplane sections of projectivized tangent bundle of projective space
by: Bansal, Ashima, et al.
Published: (2026)
by: Bansal, Ashima, et al.
Published: (2026)
Extremal Contraction of Projective Bundles
by: Bansal, Ashima, et al.
Published: (2024)
by: Bansal, Ashima, et al.
Published: (2024)
Symmetric power of higher dimensional varieties
by: Bansal, Ashima, et al.
Published: (2025)
by: Bansal, Ashima, et al.
Published: (2025)
When Does Predictive Inverse Dynamics Outperform Behavior Cloning?
by: Schäfer, Lukas, et al.
Published: (2026)
by: Schäfer, Lukas, et al.
Published: (2026)
Extracting Superficial Scattering by Q‐Sensing Technique
by: Alon Tzroya, et al.
Published: (2024)
by: Alon Tzroya, et al.
Published: (2024)
Is Behavior Cloning All You Need? Understanding Horizon in Imitation Learning
by: Foster, Dylan J., et al.
Published: (2024)
by: Foster, Dylan J., et al.
Published: (2024)
When is Cat(Q) cartesian closed?
by: Stubbe, Isar, et al.
Published: (2025)
by: Stubbe, Isar, et al.
Published: (2025)
When Life Gives You Lemons, Squeeze Your Way Through: Understanding Citrus Avoidance Behaviour by Free-Ranging Dogs in India
by: Pal, Tuhin Subhra, et al.
Published: (2024)
by: Pal, Tuhin Subhra, et al.
Published: (2024)
Sceniris: A Fast Procedural Scene Generation Framework
by: Shang, Jinghuan, et al.
Published: (2025)
by: Shang, Jinghuan, et al.
Published: (2025)
The Stepanov theorem for Q-valued functions
by: De Donato, Paolo
Published: (2024)
by: De Donato, Paolo
Published: (2024)
Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency model
by: Zhang, Jing, et al.
Published: (2024)
by: Zhang, Jing, et al.
Published: (2024)
Image and Status of the Library and Information Services Field. Final Report.
by: Walters, J. Hart, Jr.
Published: (1970)
by: Walters, J. Hart, Jr.
Published: (1970)
When Life Gives You AI, Will You Turn It Into A Market for Lemons? Understanding How Information Asymmetries About AI System Capabilities Affect Market Outcomes and Adoption
by: Erlei, Alexander, et al.
Published: (2026)
by: Erlei, Alexander, et al.
Published: (2026)
Learning Efficient and Robust Language-conditioned Manipulation using Textual-Visual Relevancy and Equivariant Language Mapping
by: Jia, Mingxi, et al.
Published: (2024)
by: Jia, Mingxi, et al.
Published: (2024)
Verifiably Following Complex Robot Instructions with Foundation Models
by: Quartey, Benedict, et al.
Published: (2024)
by: Quartey, Benedict, et al.
Published: (2024)
LEGS-POMDP: Language and Gesture-Guided Object Search in Partially Observable Environments
by: He, Ivy Xiao, et al.
Published: (2026)
by: He, Ivy Xiao, et al.
Published: (2026)
Understanding Multimodal Failure in Action-Chunking Behavioral Cloning
by: Mazza, Lorenzo, et al.
Published: (2026)
by: Mazza, Lorenzo, et al.
Published: (2026)
Equivariant Offline Reinforcement Learning
by: Tangri, Arsh, et al.
Published: (2024)
by: Tangri, Arsh, et al.
Published: (2024)
Similar Items
-
Accelerating Residual Reinforcement Learning with Uncertainty Estimation
by: Dodeja, Lakshita, et al.
Published: (2025) -
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning
by: Schroeder, Philip, et al.
Published: (2026) -
You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector
by: Patil, Omkar, et al.
Published: (2026) -
Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies
by: Huang, Haojie, et al.
Published: (2024) -
On-Robot Reinforcement Learning with Goal-Contrastive Rewards
by: Biza, Ondrej, et al.
Published: (2024)