TGRL: An Algorithm for Teacher Guided Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Shenfeld, Idan, Hong, Zhang-Wei, Tamar, Aviv, Agrawal, Pulkit |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RL's Razor: Why Online Reinforcement Learning Forgets Less
by: Shenfeld, Idan, et al.
Published: (2025)
by: Shenfeld, Idan, et al.
Published: (2025)
Self-Distillation Enables Continual Learning
by: Shenfeld, Idan, et al.
Published: (2026)
by: Shenfeld, Idan, et al.
Published: (2026)
JUICER: Data-Efficient Imitation Learning for Robotic Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
Language Model Personalization via Reward Factorization
by: Shenfeld, Idan, et al.
Published: (2025)
by: Shenfeld, Idan, et al.
Published: (2025)
From Imitation to Refinement -- Residual RL for Precise Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
Value Augmented Sampling for Language Model Alignment and Personalization
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
Test-Time Regret Minimization in Meta Reinforcement Learning
by: Mutti, Mirco, et al.
Published: (2024)
by: Mutti, Mirco, et al.
Published: (2024)
Random Latent Exploration for Deep Reinforcement Learning
by: Mahankali, Srinath, et al.
Published: (2024)
by: Mahankali, Srinath, et al.
Published: (2024)
Curiosity-driven Red-teaming for Large Language Models
by: Hong, Zhang-Wei, et al.
Published: (2024)
by: Hong, Zhang-Wei, et al.
Published: (2024)
Meta Reinforcement Learning with Finite Training Tasks -- a Density Estimation Approach
by: Rimon, Zohar, et al.
Published: (2022)
by: Rimon, Zohar, et al.
Published: (2022)
Vector Policy Optimization: Training for Diversity Improves Test-Time Search
by: Bahlous-Boldi, Ryan, et al.
Published: (2026)
by: Bahlous-Boldi, Ryan, et al.
Published: (2026)
Entity-Centric Reinforcement Learning for Object Manipulation from Pixels
by: Haramati, Dan, et al.
Published: (2024)
by: Haramati, Dan, et al.
Published: (2024)
MAMBA: an Effective World Model Approach for Meta-Reinforcement Learning
by: Rimon, Zohar, et al.
Published: (2024)
by: Rimon, Zohar, et al.
Published: (2024)
Going Beyond Heuristics by Imposing Policy Improvement as a Constraint
by: Lee, Chi-Chang, et al.
Published: (2025)
by: Lee, Chi-Chang, et al.
Published: (2025)
KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
by: Aminian, Gholamali, et al.
Published: (2025)
by: Aminian, Gholamali, et al.
Published: (2025)
DDLP: Unsupervised Object-Centric Video Prediction with Deep Dynamic Latent Particles
by: Daniel, Tal, et al.
Published: (2023)
by: Daniel, Tal, et al.
Published: (2023)
FAST-Q: Fast-track Exploration with Adversarially Balanced State Representations for Counterfactual Action Estimation in Offline Reinforcement Learning
by: Agrawal, Pulkit, et al.
Published: (2025)
by: Agrawal, Pulkit, et al.
Published: (2025)
A Classification View on Meta Learning Bandits
by: Mutti, Mirco, et al.
Published: (2025)
by: Mutti, Mirco, et al.
Published: (2025)
Few-Shot Task Learning through Inverse Generative Modeling
by: Netanyahu, Aviv, et al.
Published: (2024)
by: Netanyahu, Aviv, et al.
Published: (2024)
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
by: Damani, Mehul, et al.
Published: (2024)
by: Damani, Mehul, et al.
Published: (2024)
Multi Task Inverse Reinforcement Learning for Common Sense Reward
by: Glazer, Neta, et al.
Published: (2024)
by: Glazer, Neta, et al.
Published: (2024)
Hierarchical Entity-centric Reinforcement Learning with Factored Subgoal Diffusion
by: Haramati, Dan, et al.
Published: (2026)
by: Haramati, Dan, et al.
Published: (2026)
RoboArm-NMP: a Learning Environment for Neural Motion Planning
by: Jurgenson, Tom, et al.
Published: (2024)
by: Jurgenson, Tom, et al.
Published: (2024)
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
by: Aminian, Gholamali, et al.
Published: (2025)
by: Aminian, Gholamali, et al.
Published: (2025)
Collective Model Intelligence Requires Compatible Specialization
by: Pari, Jyothish, et al.
Published: (2024)
by: Pari, Jyothish, et al.
Published: (2024)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
Reinforcement Learning via Self-Distillation
by: Hübotter, Jonas, et al.
Published: (2026)
by: Hübotter, Jonas, et al.
Published: (2026)
Toward Artificial Palpation: Representation Learning of Touch on Soft Bodies
by: Rimon, Zohar, et al.
Published: (2025)
by: Rimon, Zohar, et al.
Published: (2025)
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning
by: Reuss, Moritz, et al.
Published: (2024)
by: Reuss, Moritz, et al.
Published: (2024)
ROER: Regularized Optimal Experience Replay
by: Li, Changling, et al.
Published: (2024)
by: Li, Changling, et al.
Published: (2024)
Explore to Generalize in Zero-Shot RL
by: Zisselman, Ev, et al.
Published: (2023)
by: Zisselman, Ev, et al.
Published: (2023)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
by: Hilel, Almog, et al.
Published: (2025)
by: Hilel, Almog, et al.
Published: (2025)
Hovering Flight of Soft-Actuated Insect-Scale Micro Aerial Vehicles using Deep Reinforcement Learning
by: Hsiao, Yi-Hsuan, et al.
Published: (2025)
by: Hsiao, Yi-Hsuan, et al.
Published: (2025)
What Happens During the Loss Plateau? Understanding Abrupt Learning in Transformers
by: Gopalani, Pulkit, et al.
Published: (2025)
by: Gopalani, Pulkit, et al.
Published: (2025)
Known Unknowns: Out-of-Distribution Property Prediction in Materials and Molecules
by: Segal, Nofit, et al.
Published: (2025)
by: Segal, Nofit, et al.
Published: (2025)
Privacy amplification by random allocation
by: Feldman, Vitaly, et al.
Published: (2025)
by: Feldman, Vitaly, et al.
Published: (2025)
Generalization in the Face of Adaptivity: A Bayesian Perspective
by: Shenfeld, Moshe, et al.
Published: (2021)
by: Shenfeld, Moshe, et al.
Published: (2021)
Efficient privacy loss accounting for subsampling and random allocation
by: Feldman, Vitaly, et al.
Published: (2026)
by: Feldman, Vitaly, et al.
Published: (2026)
Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models
by: Francis-Meretzki, Shelly, et al.
Published: (2026)
by: Francis-Meretzki, Shelly, et al.
Published: (2026)
Similar Items
-
RL's Razor: Why Online Reinforcement Learning Forgets Less
by: Shenfeld, Idan, et al.
Published: (2025) -
Self-Distillation Enables Continual Learning
by: Shenfeld, Idan, et al.
Published: (2026) -
JUICER: Data-Efficient Imitation Learning for Robotic Assembly
by: Ankile, Lars, et al.
Published: (2024) -
Language Model Personalization via Reward Factorization
by: Shenfeld, Idan, et al.
Published: (2025) -
From Imitation to Refinement -- Residual RL for Precise Assembly
by: Ankile, Lars, et al.
Published: (2024)