Intentionally-underestimated Value Function at Terminal State for Temporal-difference Learning with Mis-designed Reward
Fuente:
arXiv
Saved in:
| Main Author: | Kobayashi, Taisuke |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Consolidated Adaptive T-soft Update for Deep Reinforcement Learning
by: Kobayashi, Taisuke
Published: (2022)
by: Kobayashi, Taisuke
Published: (2022)
CubeDAgger: Interactive Imitation Learning for Dynamic Systems with Efficient yet Low-risk Interaction
by: Kobayashi, Taisuke
Published: (2025)
by: Kobayashi, Taisuke
Published: (2025)
Weber-Fechner Law in Temporal Difference learning derived from Control as Inference
by: Takahashi, Keiichiro, et al.
Published: (2024)
by: Takahashi, Keiichiro, et al.
Published: (2024)
Variational Adaptive Noise and Dropout towards Stable Recurrent Neural Networks
by: Kobayashi, Taisuke, et al.
Published: (2025)
by: Kobayashi, Taisuke, et al.
Published: (2025)
Towards Autonomous Driving of Personal Mobility with Small and Noisy Dataset using Tsallis-statistics-based Behavioral Cloning
by: Kobayashi, Taisuke, et al.
Published: (2021)
by: Kobayashi, Taisuke, et al.
Published: (2021)
Design of Restricted Normalizing Flow towards Arbitrary Stochastic Policy with Computational Efficiency
by: Kobayashi, Taisuke, et al.
Published: (2024)
by: Kobayashi, Taisuke, et al.
Published: (2024)
Data-driven Probabilistic Trajectory Learning with High Temporal Resolution in Terminal Airspace
by: Xiang, Jun, et al.
Published: (2024)
by: Xiang, Jun, et al.
Published: (2024)
Curriculum Reinforcement Learning for Complex Reward Functions
by: Freitag, Kilian, et al.
Published: (2024)
by: Freitag, Kilian, et al.
Published: (2024)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
by: Ishihara, Yu, et al.
Published: (2025)
by: Ishihara, Yu, et al.
Published: (2025)
Robot Policy Learning with Temporal Optimal Transport Reward
by: Fu, Yuwei, et al.
Published: (2024)
by: Fu, Yuwei, et al.
Published: (2024)
Pseudo-Quantized Actor-Critic Algorithm for Robustness to Noisy Temporal Difference Error
by: Kobayashi, Taisuke
Published: (2026)
by: Kobayashi, Taisuke
Published: (2026)
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
by: Liu, Yuyang, et al.
Published: (2025)
by: Liu, Yuyang, et al.
Published: (2025)
CaT: Constraints as Terminations for Legged Locomotion Reinforcement Learning
by: Chane-Sane, Elliot, et al.
Published: (2024)
by: Chane-Sane, Elliot, et al.
Published: (2024)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
A Generalized Acquisition Function for Preference-based Reward Learning
by: Ellis, Evan, et al.
Published: (2024)
by: Ellis, Evan, et al.
Published: (2024)
LiRA: Light-Robust Adversary for Model-based Reinforcement Learning in Real World
by: Kobayashi, Taisuke
Published: (2024)
by: Kobayashi, Taisuke
Published: (2024)
A Review of Reward Functions for Reinforcement Learning in the context of Autonomous Driving
by: Abouelazm, Ahmed, et al.
Published: (2024)
by: Abouelazm, Ahmed, et al.
Published: (2024)
Robotic Skill Diversification via Active Mutation of Reward Functions in Reinforcement Learning During a Liquid Pouring Task
by: van Buuren, Jannick, et al.
Published: (2025)
by: van Buuren, Jannick, et al.
Published: (2025)
On-Robot Reinforcement Learning with Goal-Contrastive Rewards
by: Biza, Ondrej, et al.
Published: (2024)
by: Biza, Ondrej, et al.
Published: (2024)
DROP: Distributional and Regular Optimism and Pessimism for Reinforcement Learning
by: Kobayashi, Taisuke
Published: (2024)
by: Kobayashi, Taisuke
Published: (2024)
A Heuristic Approach for Performance Tuning in RL-based Quadrotor Control via Reward Design and Termination Conditions
by: Suarez, Fausto Mauricio Lagos, et al.
Published: (2026)
by: Suarez, Fausto Mauricio Lagos, et al.
Published: (2026)
Mollified Value Learning
by: Viswanath, Hrishikesh, et al.
Published: (2026)
by: Viswanath, Hrishikesh, et al.
Published: (2026)
Revisiting Sparse Rewards for Goal-Reaching Reinforcement Learning
by: Vasan, Gautham, et al.
Published: (2024)
by: Vasan, Gautham, et al.
Published: (2024)
Goal-Conditioned Terminal Value Estimation for Real-time and Multi-task Model Predictive Control
by: Morita, Mitsuki, et al.
Published: (2024)
by: Morita, Mitsuki, et al.
Published: (2024)
Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation
by: Kobayashi, Masato, et al.
Published: (2025)
by: Kobayashi, Masato, et al.
Published: (2025)
Safe Value Functions
by: Massiani, Pierre-François, et al.
Published: (2021)
by: Massiani, Pierre-François, et al.
Published: (2021)
Generalization in Deep Reinforcement Learning for Robotic Navigation by Reward Shaping
by: Miranda, Victor R. F., et al.
Published: (2022)
by: Miranda, Victor R. F., et al.
Published: (2022)
Rewarding DINO: Predicting Dense Rewards with Vision Foundation Models
by: Krack, Pierre, et al.
Published: (2026)
by: Krack, Pierre, et al.
Published: (2026)
Hierarchical Reinforcement Learning Framework for Adaptive Walking Control Using General Value Functions of Lower-Limb Sensor Signals
by: Jones, Sonny T., et al.
Published: (2025)
by: Jones, Sonny T., et al.
Published: (2025)
Video2Reward: Generating Reward Function from Videos for Legged Robot Behavior Learning
by: Zeng, Runhao, et al.
Published: (2024)
by: Zeng, Runhao, et al.
Published: (2024)
Online Intention Prediction via Control-Informed Learning
by: Zhou, Tianyu, et al.
Published: (2026)
by: Zhou, Tianyu, et al.
Published: (2026)
The Dark Side of Rich Rewards: Understanding and Mitigating Noise in VLM Rewards
by: Huang, Sukai, et al.
Published: (2024)
by: Huang, Sukai, et al.
Published: (2024)
ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation
by: Feng, Youhe, et al.
Published: (2026)
by: Feng, Youhe, et al.
Published: (2026)
ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation
by: Tang, Nan, et al.
Published: (2025)
by: Tang, Nan, et al.
Published: (2025)
FLoRA: Sample-Efficient Preference-based RL via Low-Rank Style Adaptation of Reward Functions
by: Marta, Daniel, et al.
Published: (2025)
by: Marta, Daniel, et al.
Published: (2025)
Safety-Critical Traffic Simulation with Adversarial Transfer of Driving Intentions
by: Huang, Zherui, et al.
Published: (2025)
by: Huang, Zherui, et al.
Published: (2025)
Value Explicit Pretraining for Learning Transferable Representations
by: Lekkala, Kiran, et al.
Published: (2023)
by: Lekkala, Kiran, et al.
Published: (2023)
Learning Emergent Gaits with Decentralized Phase Oscillators: on the role of Observations, Rewards, and Feedback
by: Zhang, Jenny, et al.
Published: (2024)
by: Zhang, Jenny, et al.
Published: (2024)
Average-Reward Maximum Entropy Reinforcement Learning for Underactuated Double Pendulum Tasks
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
by: Chakraborty, Souradip, et al.
Published: (2023)
by: Chakraborty, Souradip, et al.
Published: (2023)
Similar Items
-
Consolidated Adaptive T-soft Update for Deep Reinforcement Learning
by: Kobayashi, Taisuke
Published: (2022) -
CubeDAgger: Interactive Imitation Learning for Dynamic Systems with Efficient yet Low-risk Interaction
by: Kobayashi, Taisuke
Published: (2025) -
Weber-Fechner Law in Temporal Difference learning derived from Control as Inference
by: Takahashi, Keiichiro, et al.
Published: (2024) -
Variational Adaptive Noise and Dropout towards Stable Recurrent Neural Networks
by: Kobayashi, Taisuke, et al.
Published: (2025) -
Towards Autonomous Driving of Personal Mobility with Small and Noisy Dataset using Tsallis-statistics-based Behavioral Cloning
by: Kobayashi, Taisuke, et al.
Published: (2021)