Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yifan, Grover, Sachin, Mistiri, Mohamed El, Kalirathnam, Kamalesh, Kerhalkar, Pratyush, Mishra, Swaroop, Kumar, Neelesh, Gaurav, Sanket, Aran, Oya, Amor, Heni Ben
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908677671223296
author Zhou, Yifan
Grover, Sachin
Mistiri, Mohamed El
Kalirathnam, Kamalesh
Kerhalkar, Pratyush
Mishra, Swaroop
Kumar, Neelesh
Gaurav, Sanket
Aran, Oya
Amor, Heni Ben
author_facet Zhou, Yifan
Grover, Sachin
Mistiri, Mohamed El
Kalirathnam, Kamalesh
Kerhalkar, Pratyush
Mishra, Swaroop
Kumar, Neelesh
Gaurav, Sanket
Aran, Oya
Amor, Heni Ben
contents Reinforcement Learning (RL) traditionally relies on scalar reward signals, limiting its ability to leverage the rich semantic knowledge often available in real-world tasks. In contrast, humans learn efficiently by combining numerical feedback with language, prior knowledge, and common sense. We introduce Prompted Policy Search (ProPS), a novel RL method that unifies numerical and linguistic reasoning within a single framework. Unlike prior work that augment existing RL components with language, ProPS places a large language model (LLM) at the center of the policy optimization loop-directly proposing policy updates based on both reward feedback and natural language input. We show that LLMs can perform numerical optimization in-context, and that incorporating semantic signals, such as goals, domain knowledge, and strategy hints can lead to more informed exploration and sample-efficient learning. ProPS is evaluated across fifteen Gymnasium tasks, spanning classic control, Atari games, and MuJoCo environments, and compared to seven widely-adopted RL algorithms (e.g., PPO, SAC, TRPO). It outperforms all baselines on eight out of fifteen tasks and demonstrates substantial gains when provided with domain knowledge. These results highlight the potential of unifying semantics and numerics for transparent, generalizable, and human-aligned RL.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21928
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
Zhou, Yifan
Grover, Sachin
Mistiri, Mohamed El
Kalirathnam, Kamalesh
Kerhalkar, Pratyush
Mishra, Swaroop
Kumar, Neelesh
Gaurav, Sanket
Aran, Oya
Amor, Heni Ben
Machine Learning
Artificial Intelligence
Reinforcement Learning (RL) traditionally relies on scalar reward signals, limiting its ability to leverage the rich semantic knowledge often available in real-world tasks. In contrast, humans learn efficiently by combining numerical feedback with language, prior knowledge, and common sense. We introduce Prompted Policy Search (ProPS), a novel RL method that unifies numerical and linguistic reasoning within a single framework. Unlike prior work that augment existing RL components with language, ProPS places a large language model (LLM) at the center of the policy optimization loop-directly proposing policy updates based on both reward feedback and natural language input. We show that LLMs can perform numerical optimization in-context, and that incorporating semantic signals, such as goals, domain knowledge, and strategy hints can lead to more informed exploration and sample-efficient learning. ProPS is evaluated across fifteen Gymnasium tasks, spanning classic control, Atari games, and MuJoCo environments, and compared to seven widely-adopted RL algorithms (e.g., PPO, SAC, TRPO). It outperforms all baselines on eight out of fifteen tasks and demonstrates substantial gains when provided with domain knowledge. These results highlight the potential of unifying semantics and numerics for transparent, generalizable, and human-aligned RL.
title Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.21928