Natural Language Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Xidong, Liu, Bo, Song, Yan, Fu, Haotian, Wan, Ziyu, Koushik, Girish A., Hu, Zhiyuan, Yang, Mengyue, Wen, Ying, Wang, Jun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909625315491840
author Feng, Xidong
Liu, Bo
Song, Yan
Fu, Haotian
Wan, Ziyu
Koushik, Girish A.
Hu, Zhiyuan
Yang, Mengyue
Wen, Ying
Wang, Jun
author_facet Feng, Xidong
Liu, Bo
Song, Yan
Fu, Haotian
Wan, Ziyu
Koushik, Girish A.
Hu, Zhiyuan
Yang, Mengyue
Wen, Ying
Wang, Jun
contents Artificial intelligence progresses towards the "Era of Experience," where agents are expected to learn from continuous, grounded interaction. We argue that traditional Reinforcement Learning (RL), which typically represents value as a scalar, can restrict agent's deep understanding of environments and hinders the active, deliberative learning crucial for navigating this new paradigm. To address the issue, we introduce Natural Language Reinforcement Learning (NLRL), a framework that extends RL principles into natural language counterparts. Central to NLRL is the Language Value Function (LVF), which redefines value as an interpretable linguistic narrative articulating the rationale behind an evaluation. NLRL further extends this concept to core RL components, including policy, the Bellman equation, and policy iteration. Leveraging recent advancements in Large Language Models (LLMs), NLRL can be practically implemented to achieve RL-like policy and value training through unsupervised environment interactions. Experiments over 4 multi-step agentic tasks demonstrate NLRL's effectiveness, efficiency, and its potential to foster deeper understanding and more active learning strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14251
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Natural Language Reinforcement Learning
Feng, Xidong
Liu, Bo
Song, Yan
Fu, Haotian
Wan, Ziyu
Koushik, Girish A.
Hu, Zhiyuan
Yang, Mengyue
Wen, Ying
Wang, Jun
Machine Learning
Artificial Intelligence
Computation and Language
Artificial intelligence progresses towards the "Era of Experience," where agents are expected to learn from continuous, grounded interaction. We argue that traditional Reinforcement Learning (RL), which typically represents value as a scalar, can restrict agent's deep understanding of environments and hinders the active, deliberative learning crucial for navigating this new paradigm. To address the issue, we introduce Natural Language Reinforcement Learning (NLRL), a framework that extends RL principles into natural language counterparts. Central to NLRL is the Language Value Function (LVF), which redefines value as an interpretable linguistic narrative articulating the rationale behind an evaluation. NLRL further extends this concept to core RL components, including policy, the Bellman equation, and policy iteration. Leveraging recent advancements in Large Language Models (LLMs), NLRL can be practically implemented to achieve RL-like policy and value training through unsupervised environment interactions. Experiments over 4 multi-step agentic tasks demonstrate NLRL's effectiveness, efficiency, and its potential to foster deeper understanding and more active learning strategies.
title Natural Language Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.14251