STO-RL: Offline RL under Sparse Rewards via LLM-Guided Subgoal Temporal Order
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Chengyang, Pan, Yuxin, Xiong, Hui, Chen, Yize |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SEAL: SEmantic-Augmented Imitation Learning via Language Model
by: Gu, Chengyang, et al.
Published: (2024)
by: Gu, Chengyang, et al.
Published: (2024)
A Pontryagin Method of Model-based Reinforcement Learning via Hamiltonian Actor-Critic
by: Gu, Chengyang, et al.
Published: (2026)
by: Gu, Chengyang, et al.
Published: (2026)
Learning and Optimization for Price-based Demand Response of Electric Vehicle Charging
by: Gu, Chengyang, et al.
Published: (2024)
by: Gu, Chengyang, et al.
Published: (2024)
Pontryagin Optimal Control via Neural Networks
by: Gu, Chengyang, et al.
Published: (2022)
by: Gu, Chengyang, et al.
Published: (2022)
Learning Hidden Subgoals under Temporal Ordering Constraints in Reinforcement Learning
by: Xu, Duo, et al.
Published: (2024)
by: Xu, Duo, et al.
Published: (2024)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Enhanced-FQL($λ$), an Efficient and Interpretable RL with novel Fuzzy Eligibility Traces and Segmented Experience Replay
by: Jalaeian-Farimani, Mohsen, et al.
Published: (2026)
by: Jalaeian-Farimani, Mohsen, et al.
Published: (2026)
Hereditary Geometric Meta-RL: Nonlocal Generalization via Task Symmetries
by: Nitschke, Paul, et al.
Published: (2026)
by: Nitschke, Paul, et al.
Published: (2026)
RL for Mitigating Cascading Failures: Targeted Exploration via Sensitivity Factors
by: Dwivedi, Anmol, et al.
Published: (2024)
by: Dwivedi, Anmol, et al.
Published: (2024)
AutoRL Hyperparameter Landscapes
by: Mohan, Aditya, et al.
Published: (2023)
by: Mohan, Aditya, et al.
Published: (2023)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
Unifying Goal-Conditioned RL and Unsupervised Skill Learning via Control-Maximization
by: Modirshanechi, Alireza, et al.
Published: (2026)
by: Modirshanechi, Alireza, et al.
Published: (2026)
Multi-Agent Path Finding via Offline RL and LLM Collaboration
by: Atasever, Merve, et al.
Published: (2025)
by: Atasever, Merve, et al.
Published: (2025)
Visual CPG-RL: Learning Central Pattern Generators for Visually-Guided Quadruped Locomotion
by: Bellegarda, Guillaume, et al.
Published: (2022)
by: Bellegarda, Guillaume, et al.
Published: (2022)
Blackout Mitigation via Physics-guided RL
by: Dwivedi, Anmol, et al.
Published: (2024)
by: Dwivedi, Anmol, et al.
Published: (2024)
MAGICS: Adversarial RL with Minimax Actors Guided by Implicit Critic Stackelberg for Convergent Neural Synthesis of Robot Safety
by: Wang, Justin, et al.
Published: (2024)
by: Wang, Justin, et al.
Published: (2024)
Align and Filter: Improving Performance in Asynchronous On-Policy RL
by: Honari, Homayoun, et al.
Published: (2026)
by: Honari, Homayoun, et al.
Published: (2026)
Mining--Gym: A Configurable RL Benchmarking Environment for Truck Dispatch Scheduling
by: Banerjee, Chayan, et al.
Published: (2025)
by: Banerjee, Chayan, et al.
Published: (2025)
CBF-RL: Safety Filtering Reinforcement Learning in Training with Control Barrier Functions
by: Yang, Lizhi, et al.
Published: (2025)
by: Yang, Lizhi, et al.
Published: (2025)
Analytical Lyapunov Function Discovery: An RL-based Generative Approach
by: Zou, Haohan, et al.
Published: (2025)
by: Zou, Haohan, et al.
Published: (2025)
EcoFair-CH-MARL: Scalable Constrained Hierarchical Multi-Agent RL with Real-Time Emission Budgets and Fairness Guarantees
by: Alqithami, Saad
Published: (2026)
by: Alqithami, Saad
Published: (2026)
Two-Stage Active Distribution Network Voltage Control via LLM-RL Collaboration: A Hybrid Knowledge-Data-Driven Approach
by: Yang, Xu, et al.
Published: (2026)
by: Yang, Xu, et al.
Published: (2026)
OMG-RL:Offline Model-based Guided Reward Learning for Heparin Treatment
by: Lim, Yooseok, et al.
Published: (2024)
by: Lim, Yooseok, et al.
Published: (2024)
Multi-Task Vehicle Routing Solver via Mixture of Specialized Experts under State-Decomposable MDP
by: Pan, Yuxin, et al.
Published: (2025)
by: Pan, Yuxin, et al.
Published: (2025)
ReLMXEL: Adaptive RL-Based Memory Controller with Explainable Energy and Latency Optimization
by: Sai, Panuganti Chirag, et al.
Published: (2026)
by: Sai, Panuganti Chirag, et al.
Published: (2026)
Digital Twin Synchronization: Bridging the Sim-RL Agent to a Real-Time Robotic Additive Manufacturing Control
by: Ali, Matsive, et al.
Published: (2025)
by: Ali, Matsive, et al.
Published: (2025)
RL-Based Method for Benchmarking the Adversarial Resilience and Robustness of Deep Reinforcement Learning Policies
by: Behzadan, Vahid, et al.
Published: (2019)
by: Behzadan, Vahid, et al.
Published: (2019)
Efficient On-policy Visual-RL via Stochastic Decoupled Policy Gradient
by: You, Haoxiang, et al.
Published: (2026)
by: You, Haoxiang, et al.
Published: (2026)
Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction
by: Durkin, Alex, et al.
Published: (2025)
by: Durkin, Alex, et al.
Published: (2025)
Stabilizing Policy Gradient Methods via Reward Profiling
by: Ahmed, Shihab, et al.
Published: (2025)
by: Ahmed, Shihab, et al.
Published: (2025)
Large Language Model Guided Incentive Aware Reward Design for Cooperative Multi-Agent Reinforcement Learning
by: Urgun, Dogan, et al.
Published: (2026)
by: Urgun, Dogan, et al.
Published: (2026)
Learning from Imperfect Demonstrations via Temporal Behavior Tree-Guided Trajectory Repair
by: Puranic, Aniruddh G., et al.
Published: (2026)
by: Puranic, Aniruddh G., et al.
Published: (2026)
Online Location Planning for AI-Defined Vehicles: Optimizing Joint Tasks of Order Serving and Spatio-Temporal Heterogeneous Model Fine-Tuning
by: Zheng, Bokeng, et al.
Published: (2025)
by: Zheng, Bokeng, et al.
Published: (2025)
Bregman Centroid Guided Cross-Entropy Method
by: Gu, Yuliang, et al.
Published: (2025)
by: Gu, Yuliang, et al.
Published: (2025)
Mutual Information as Intrinsic Reward of Reinforcement Learning Agents for On-demand Ride Pooling
by: Zhang, Xianjie, et al.
Published: (2023)
by: Zhang, Xianjie, et al.
Published: (2023)
Offline Reinforcement Learning and Sequence Modeling for Downlink Link Adaptation
by: Peri, Samuele, et al.
Published: (2024)
by: Peri, Samuele, et al.
Published: (2024)
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
by: Wang, Taiyi, et al.
Published: (2024)
by: Wang, Taiyi, et al.
Published: (2024)
Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving
by: Zhang, Zhihao, et al.
Published: (2025)
by: Zhang, Zhihao, et al.
Published: (2025)
Continuously Improving Mobile Manipulation with Autonomous Real-World RL
by: Mendonca, Russell, et al.
Published: (2024)
by: Mendonca, Russell, et al.
Published: (2024)
Data Center Cooling System Optimization Using Offline Reinforcement Learning
by: Zhan, Xianyuan, et al.
Published: (2025)
by: Zhan, Xianyuan, et al.
Published: (2025)
Similar Items
-
SEAL: SEmantic-Augmented Imitation Learning via Language Model
by: Gu, Chengyang, et al.
Published: (2024) -
A Pontryagin Method of Model-based Reinforcement Learning via Hamiltonian Actor-Critic
by: Gu, Chengyang, et al.
Published: (2026) -
Learning and Optimization for Price-based Demand Response of Electric Vehicle Charging
by: Gu, Chengyang, et al.
Published: (2024) -
Pontryagin Optimal Control via Neural Networks
by: Gu, Chengyang, et al.
Published: (2022) -
Learning Hidden Subgoals under Temporal Ordering Constraints in Reinforcement Learning
by: Xu, Duo, et al.
Published: (2024)