Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Shuze Daniel, Chen, Claire, Xiao, Jiabao Sean, Lei, Lei, Zhang, Yuheng, Yue, Yisong, Simchi-Levi, David
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908954267746304
author Liu, Shuze Daniel
Chen, Claire
Xiao, Jiabao Sean
Lei, Lei
Zhang, Yuheng
Yue, Yisong
Simchi-Levi, David
author_facet Liu, Shuze Daniel
Chen, Claire
Xiao, Jiabao Sean
Lei, Lei
Zhang, Yuheng
Yue, Yisong
Simchi-Levi, David
contents The recent advancement of Large Language Models (LLMs) has established their potential as autonomous interactive agents. However, they often struggle in strategic games of incomplete information, such as bilateral price negotiation. In this paper, we investigate if Reinforcement Learning from Verifiable Rewards (RLVR) can effectively teach LLMs to negotiate. Specifically, we explore the strategic behaviors that emerge during the learning process. We introduce a framework that trains a mid-sized buyer agent against a regulated LLM seller across a wide distribution of real-world products. By grounding reward signals directly in the maximization of economic surplus and strict adherence to private budget constraints, we reveal a novel four-phase strategic evolution. The agent progresses from naive bargaining to using aggressive starting prices, moves through a phase of deadlock, and ultimately develops sophisticated persuasive skills. Our results demonstrate that this verifiable training allows a 30B agent to significantly outperform frontier models over ten times its size in extracting surplus. Furthermore, the trained agent generalizes robustly to stronger counterparties unseen during training and remains effective even when facing hostile, adversarial seller personas.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09855
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards
Liu, Shuze Daniel
Chen, Claire
Xiao, Jiabao Sean
Lei, Lei
Zhang, Yuheng
Yue, Yisong
Simchi-Levi, David
Artificial Intelligence
Computation and Language
Computer Science and Game Theory
General Economics
Economics
The recent advancement of Large Language Models (LLMs) has established their potential as autonomous interactive agents. However, they often struggle in strategic games of incomplete information, such as bilateral price negotiation. In this paper, we investigate if Reinforcement Learning from Verifiable Rewards (RLVR) can effectively teach LLMs to negotiate. Specifically, we explore the strategic behaviors that emerge during the learning process. We introduce a framework that trains a mid-sized buyer agent against a regulated LLM seller across a wide distribution of real-world products. By grounding reward signals directly in the maximization of economic surplus and strict adherence to private budget constraints, we reveal a novel four-phase strategic evolution. The agent progresses from naive bargaining to using aggressive starting prices, moves through a phase of deadlock, and ultimately develops sophisticated persuasive skills. Our results demonstrate that this verifiable training allows a 30B agent to significantly outperform frontier models over ten times its size in extracting surplus. Furthermore, the trained agent generalizes robustly to stronger counterparties unseen during training and remains effective even when facing hostile, adversarial seller personas.
title Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards
topic Artificial Intelligence
Computation and Language
Computer Science and Game Theory
General Economics
Economics
url https://arxiv.org/abs/2604.09855