TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Haotian, Wang, Pengcheng, Schneider, Jeff, Shi, Guanya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoVO-MPC: Theoretical Analysis of Sampling-based MPC and Optimal Covariance Design
by: Yi, Zeji, et al.
Published: (2024)
by: Yi, Zeji, et al.
Published: (2024)
TD-MPC-Opt: Distilling Model-Based Multi-Task Reinforcement Learning Agents
by: Kuzmenko, Dmytro, et al.
Published: (2025)
by: Kuzmenko, Dmytro, et al.
Published: (2025)
DADP: Domain Adaptive Diffusion Policy
by: Wang, Pengcheng, et al.
Published: (2026)
by: Wang, Pengcheng, et al.
Published: (2026)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
by: Ankile, Lars, et al.
Published: (2025)
by: Ankile, Lars, et al.
Published: (2025)
Safe Deep Policy Adaptation
by: Xiao, Wenli, et al.
Published: (2023)
by: Xiao, Wenli, et al.
Published: (2023)
TD-MPC2: Scalable, Robust World Models for Continuous Control
by: Hansen, Nicklas, et al.
Published: (2023)
by: Hansen, Nicklas, et al.
Published: (2023)
Latent Policy Steering with Embodiment-Agnostic Pretrained World Models
by: Wang, Yiqi, et al.
Published: (2025)
by: Wang, Yiqi, et al.
Published: (2025)
Bridging Adaptivity and Safety: Learning Agile Collision-Free Locomotion Across Varied Physics
by: Zhong, Yichao, et al.
Published: (2025)
by: Zhong, Yichao, et al.
Published: (2025)
Toward Scalable Multirobot Control: Fast Policy Learning in Distributed MPC
by: Zhang, Xinglong, et al.
Published: (2024)
by: Zhang, Xinglong, et al.
Published: (2024)
TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion
by: Nguyen, Khang, et al.
Published: (2025)
by: Nguyen, Khang, et al.
Published: (2025)
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning
by: Venugopal, Aravind, et al.
Published: (2026)
by: Venugopal, Aravind, et al.
Published: (2026)
ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
by: Zhao, Siheng, et al.
Published: (2025)
by: Zhao, Siheng, et al.
Published: (2025)
Residual Neural Terminal Constraint for MPC-based Collision Avoidance in Dynamic Environments
by: Derajić, Bojan, et al.
Published: (2025)
by: Derajić, Bojan, et al.
Published: (2025)
Diffusion Policy with Bayesian Expert Selection for Active Multi-Target Tracking
by: Xiang, Haotian, et al.
Published: (2026)
by: Xiang, Haotian, et al.
Published: (2026)
Exterior Penalty Policy Optimization with Penalty Metric Network under Constraints
by: Gao, Shiqing, et al.
Published: (2024)
by: Gao, Shiqing, et al.
Published: (2024)
Dataset Clustering for Improved Offline Policy Learning
by: Wang, Qiang, et al.
Published: (2024)
by: Wang, Qiang, et al.
Published: (2024)
EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control
by: Evers, Thomas, et al.
Published: (2026)
by: Evers, Thomas, et al.
Published: (2026)
Model-Based Diffusion for Trajectory Optimization
by: Pan, Chaoyi, et al.
Published: (2024)
by: Pan, Chaoyi, et al.
Published: (2024)
Assigning Credit with Partial Reward Decoupling in Multi-Agent Proximal Policy Optimization
by: Kapoor, Aditya, et al.
Published: (2024)
by: Kapoor, Aditya, et al.
Published: (2024)
Shielded Reinforcement Learning Under Dynamic Temporal Logic Constraints
by: Yüksel, Sadık Bera, et al.
Published: (2026)
by: Yüksel, Sadık Bera, et al.
Published: (2026)
Planning with Adaptive World Models for Autonomous Driving
by: Vasudevan, Arun Balajee, et al.
Published: (2024)
by: Vasudevan, Arun Balajee, et al.
Published: (2024)
Learning Sim-to-Real Humanoid Locomotion in 15 Minutes
by: Seo, Younggyo, et al.
Published: (2025)
by: Seo, Younggyo, et al.
Published: (2025)
Collision Probability Distribution Estimation via Temporal Difference Learning
by: Steinecker, Thomas, et al.
Published: (2024)
by: Steinecker, Thomas, et al.
Published: (2024)
Performance-driven Constrained Optimal Auto-Tuner for MPC
by: Puigjaner, Albert Gassol, et al.
Published: (2025)
by: Puigjaner, Albert Gassol, et al.
Published: (2025)
Implementing TD3 to train a Neural Network to fly a Quadcopter through an FPV Gate
by: Thomas, Patrick, et al.
Published: (2024)
by: Thomas, Patrick, et al.
Published: (2024)
Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
Accelerated Online Reinforcement Learning using Auxiliary Start State Distributions
by: Mehra, Aman, et al.
Published: (2025)
by: Mehra, Aman, et al.
Published: (2025)
Physics-informed Temporal Difference Metric Learning for Robot Motion Planning
by: Ni, Ruiqi, et al.
Published: (2025)
by: Ni, Ruiqi, et al.
Published: (2025)
LTLDoG: Satisfying Temporally-Extended Symbolic Constraints for Safe Diffusion-based Planning
by: Feng, Zeyu, et al.
Published: (2024)
by: Feng, Zeyu, et al.
Published: (2024)
SOMTP: Self-Supervised Learning-Based Optimizer for MPC-Based Safe Trajectory Planning Problems in Robotics
by: Liu, Yifan, et al.
Published: (2024)
by: Liu, Yifan, et al.
Published: (2024)
Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models
by: Francis-Meretzki, Shelly, et al.
Published: (2026)
by: Francis-Meretzki, Shelly, et al.
Published: (2026)
Weber-Fechner Law in Temporal Difference learning derived from Control as Inference
by: Takahashi, Keiichiro, et al.
Published: (2024)
by: Takahashi, Keiichiro, et al.
Published: (2024)
End-to-end RL Improves Dexterous Grasping Policies
by: Singh, Ritvik, et al.
Published: (2025)
by: Singh, Ritvik, et al.
Published: (2025)
V2X-VLM: End-to-End V2X Cooperative Autonomous Driving Through Large Vision-Language Models
by: You, Junwei, et al.
Published: (2024)
by: You, Junwei, et al.
Published: (2024)
Tractable Joint Prediction and Planning over Discrete Behavior Modes for Urban Driving
by: Villaflor, Adam, et al.
Published: (2024)
by: Villaflor, Adam, et al.
Published: (2024)
Hybrid DQN-TD3 Reinforcement Learning for Autonomous Navigation in Dynamic Environments
by: He, Xiaoyi, et al.
Published: (2025)
by: He, Xiaoyi, et al.
Published: (2025)
FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
by: Seo, Younggyo, et al.
Published: (2025)
by: Seo, Younggyo, et al.
Published: (2025)
Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition
by: Cao, Jiahang, et al.
Published: (2025)
by: Cao, Jiahang, et al.
Published: (2025)
Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies
by: Huang, Haojie, et al.
Published: (2024)
by: Huang, Haojie, et al.
Published: (2024)
WarmPrior: Straightening Flow-Matching Policies with Temporal Priors
by: Kang, Sinjae, et al.
Published: (2026)
by: Kang, Sinjae, et al.
Published: (2026)
Similar Items
-
CoVO-MPC: Theoretical Analysis of Sampling-based MPC and Optimal Covariance Design
by: Yi, Zeji, et al.
Published: (2024) -
TD-MPC-Opt: Distilling Model-Based Multi-Task Reinforcement Learning Agents
by: Kuzmenko, Dmytro, et al.
Published: (2025) -
DADP: Domain Adaptive Diffusion Policy
by: Wang, Pengcheng, et al.
Published: (2026) -
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
by: Ankile, Lars, et al.
Published: (2025) -
Safe Deep Policy Adaptation
by: Xiao, Wenli, et al.
Published: (2023)