Beyond Single-Step Updates: Reinforcement Learning of Heuristics with Limited-Horizon Search
Fuente:
arXiv
Saved in:
| Main Authors: | Hadar, Gal, Agostinelli, Forest, Shperberg, Shahaf S. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A* Search Without Expansions: Learning Heuristic Functions with Deep Q-Networks
by: Agostinelli, Forest, et al.
Published: (2021)
by: Agostinelli, Forest, et al.
Published: (2021)
The DeepXube Software Package for Solving Pathfinding Problems with Learned Heuristic Functions and Search
by: Agostinelli, Forest
Published: (2026)
by: Agostinelli, Forest
Published: (2026)
Bidirectional Bounded-Suboptimal Heuristic Search with Consistent Heuristics
by: Shperberg, Shahaf S., et al.
Published: (2025)
by: Shperberg, Shahaf S., et al.
Published: (2025)
Towards Learning Foundation Models for Heuristic Functions to Solve Pathfinding Problems
by: Khandelwal, Vedant, et al.
Published: (2024)
by: Khandelwal, Vedant, et al.
Published: (2024)
Integrating Reinforcement Learning, Action Model Learning, and Numeric Planning for Tackling Complex Tasks
by: Benyamin, Yarin, et al.
Published: (2025)
by: Benyamin, Yarin, et al.
Published: (2025)
From Kinematics to Dynamics: Learning to Refine Hybrid Plans for Physically Feasible Execution
by: Erez, Lidor, et al.
Published: (2026)
by: Erez, Lidor, et al.
Published: (2026)
On Parallel External-Memory Bidirectional Search
by: Siag, Lior, et al.
Published: (2024)
by: Siag, Lior, et al.
Published: (2024)
RAMP: Hybrid DRL for Online Learning of Numeric Action Models
by: Benyamin, Yarin, et al.
Published: (2026)
by: Benyamin, Yarin, et al.
Published: (2026)
Learning Safe Numeric Planning Action Models
by: Mordoch, Argaman, et al.
Published: (2023)
by: Mordoch, Argaman, et al.
Published: (2023)
Toward PDDL Planning Copilot
by: Benyamin, Yarin, et al.
Published: (2025)
by: Benyamin, Yarin, et al.
Published: (2025)
PDDLFuse: A Tool for Generating Diverse Planning Domains
by: Khandelwal, Vedant, et al.
Published: (2024)
by: Khandelwal, Vedant, et al.
Published: (2024)
Planning and Acting While the Clock Ticks
by: Coles, Andrew, et al.
Published: (2024)
by: Coles, Andrew, et al.
Published: (2024)
Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2024)
by: Vincent, Théo, et al.
Published: (2024)
Enhancing Metacognitive AI: Knowledge-Graph Population with Graph-Theoretic LLM Enrichment
by: Askin, Deniz, et al.
Published: (2026)
by: Askin, Deniz, et al.
Published: (2026)
Student Engagement in AI Assisted Complex Problem Solving: A Pilot Study of Human AI Rubik's Cube Collaboration
by: Vanacore, Kirk, et al.
Published: (2025)
by: Vanacore, Kirk, et al.
Published: (2025)
Subgoal-Guided Policy Heuristic Search with Learned Subgoals
by: Tuero, Jake, et al.
Published: (2025)
by: Tuero, Jake, et al.
Published: (2025)
Multiagent Reinforcement Learning for Liquidity Games
by: Vidler, Alicia, et al.
Published: (2026)
by: Vidler, Alicia, et al.
Published: (2026)
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
by: Ding, Zihan, et al.
Published: (2024)
by: Ding, Zihan, et al.
Published: (2024)
Modular Reinforcement Learning For Cooperative Swarms
by: Shtossel, Erel, et al.
Published: (2026)
by: Shtossel, Erel, et al.
Published: (2026)
Transitive Expert Error and Routing Problems in Complex AI Systems
by: Mars, Forest
Published: (2026)
by: Mars, Forest
Published: (2026)
Horizon Generalization in Reinforcement Learning
by: Myers, Vivek, et al.
Published: (2025)
by: Myers, Vivek, et al.
Published: (2025)
Is Geometry Enough? An Evaluation of Landmark-Based Gaze Estimation
by: Agostinelli, Daniele, et al.
Published: (2026)
by: Agostinelli, Daniele, et al.
Published: (2026)
Generalizing Multi-Objective Search via Objective-Aggregation Functions
by: Peer, Hadar, et al.
Published: (2025)
by: Peer, Hadar, et al.
Published: (2025)
Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities
by: Abraham, Armaan A., et al.
Published: (2026)
by: Abraham, Armaan A., et al.
Published: (2026)
Heuristic Search for Multi-Objective Probabilistic Planning
by: Chen, Dillon, et al.
Published: (2023)
by: Chen, Dillon, et al.
Published: (2023)
Gradient Boosting Reinforcement Learning
by: Fuhrer, Benjamin, et al.
Published: (2024)
by: Fuhrer, Benjamin, et al.
Published: (2024)
Heuristic Transformer: Belief Augmented In-Context Reinforcement Learning
by: Dippel, Oliver, et al.
Published: (2025)
by: Dippel, Oliver, et al.
Published: (2025)
On the Effective Horizon of Inverse Reinforcement Learning
by: Xu, Yiqing, et al.
Published: (2023)
by: Xu, Yiqing, et al.
Published: (2023)
Beyond Inference-Time Search: Reinforcement Learning Synthesizes Reusable Solvers
by: Massoudi, Soheyl, et al.
Published: (2026)
by: Massoudi, Soheyl, et al.
Published: (2026)
Combining Monte Carlo Tree Search and Heuristic Search for Weighted Vertex Coloring
by: Grelier, Cyril, et al.
Published: (2023)
by: Grelier, Cyril, et al.
Published: (2023)
Inverse Design in Distributed Circuits Using Single-Step Reinforcement Learning
by: Li, Jiayu, et al.
Published: (2025)
by: Li, Jiayu, et al.
Published: (2025)
Vanishing Bias Heuristic-guided Reinforcement Learning Algorithm
by: Li, Qinru, et al.
Published: (2023)
by: Li, Qinru, et al.
Published: (2023)
ARMS: Automatic Reward Shaping for Sparse-Reward Multi-Agent Reinforcement Learning
by: Abboud, Elie, et al.
Published: (2026)
by: Abboud, Elie, et al.
Published: (2026)
StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning
by: Zhang, Yanfei, et al.
Published: (2026)
by: Zhang, Yanfei, et al.
Published: (2026)
Rewarding What Matters: Step-by-Step Reinforcement Learning for Task-Oriented Dialogue
by: Du, Huifang, et al.
Published: (2024)
by: Du, Huifang, et al.
Published: (2024)
Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search
by: Peer, Hadar, et al.
Published: (2026)
by: Peer, Hadar, et al.
Published: (2026)
Intentional Updates for Streaming Reinforcement Learning
by: Sharifnassab, Arsalan, et al.
Published: (2026)
by: Sharifnassab, Arsalan, et al.
Published: (2026)
Planning of Heuristics: Strategic Planning on Large Language Models with Monte Carlo Tree Search for Automating Heuristic Optimization
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
Personalized Medication Planning via Direct Domain Modeling and LLM-Generated Heuristics
by: Vernik, Yonatan, et al.
Published: (2026)
by: Vernik, Yonatan, et al.
Published: (2026)
VAULT: Vigilant Adversarial Updates via LLM-Driven Retrieval-Augmented Generation for NLI
by: Kazoom, Roie, et al.
Published: (2025)
by: Kazoom, Roie, et al.
Published: (2025)
Similar Items
-
A* Search Without Expansions: Learning Heuristic Functions with Deep Q-Networks
by: Agostinelli, Forest, et al.
Published: (2021) -
The DeepXube Software Package for Solving Pathfinding Problems with Learned Heuristic Functions and Search
by: Agostinelli, Forest
Published: (2026) -
Bidirectional Bounded-Suboptimal Heuristic Search with Consistent Heuristics
by: Shperberg, Shahaf S., et al.
Published: (2025) -
Towards Learning Foundation Models for Heuristic Functions to Solve Pathfinding Problems
by: Khandelwal, Vedant, et al.
Published: (2024) -
Integrating Reinforcement Learning, Action Model Learning, and Numeric Planning for Tackling Complex Tasks
by: Benyamin, Yarin, et al.
Published: (2025)