PRACT: Optimizing Principled Reasoning and Acting of LLM Agent
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929556798046208 |
|---|---|
| author | Liu, Zhiwei Yao, Weiran Zhang, Jianguo Murthy, Rithesh Yang, Liangwei Liu, Zuxin Lan, Tian Zhu, Ming Tan, Juntao Kokane, Shirley Hoang, Thai Niebles, Juan Carlos Heinecke, Shelby Wang, Huan Savarese, Silvio Xiong, Caiming |
| author_facet | Liu, Zhiwei Yao, Weiran Zhang, Jianguo Murthy, Rithesh Yang, Liangwei Liu, Zuxin Lan, Tian Zhu, Ming Tan, Juntao Kokane, Shirley Hoang, Thai Niebles, Juan Carlos Heinecke, Shelby Wang, Huan Savarese, Silvio Xiong, Caiming |
| contents | We introduce the Principled Reasoning and Acting (PRAct) framework, a novel method for learning and enforcing action principles from trajectory data. Central to our approach is the use of text gradients from a reflection and optimization engine to derive these action principles. To adapt action principles to specific task requirements, we propose a new optimization framework, Reflective Principle Optimization (RPO). After execution, RPO employs a reflector to critique current action principles and an optimizer to update them accordingly. We develop the RPO framework under two scenarios: Reward-RPO, which uses environmental rewards for reflection, and Self-RPO, which conducts self-reflection without external rewards. Additionally, two RPO methods, RPO-Traj and RPO-Batch, is introduced to adapt to different settings. Experimental results across four environments demonstrate that the PRAct agent, leveraging the RPO framework, effectively learns and applies action principles to enhance performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_18528 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | PRACT: Optimizing Principled Reasoning and Acting of LLM Agent Liu, Zhiwei Yao, Weiran Zhang, Jianguo Murthy, Rithesh Yang, Liangwei Liu, Zuxin Lan, Tian Zhu, Ming Tan, Juntao Kokane, Shirley Hoang, Thai Niebles, Juan Carlos Heinecke, Shelby Wang, Huan Savarese, Silvio Xiong, Caiming Artificial Intelligence We introduce the Principled Reasoning and Acting (PRAct) framework, a novel method for learning and enforcing action principles from trajectory data. Central to our approach is the use of text gradients from a reflection and optimization engine to derive these action principles. To adapt action principles to specific task requirements, we propose a new optimization framework, Reflective Principle Optimization (RPO). After execution, RPO employs a reflector to critique current action principles and an optimizer to update them accordingly. We develop the RPO framework under two scenarios: Reward-RPO, which uses environmental rewards for reflection, and Self-RPO, which conducts self-reflection without external rewards. Additionally, two RPO methods, RPO-Traj and RPO-Batch, is introduced to adapt to different settings. Experimental results across four environments demonstrate that the PRAct agent, leveraging the RPO framework, effectively learns and applies action principles to enhance performance. |
| title | PRACT: Optimizing Principled Reasoning and Acting of LLM Agent |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2410.18528 |