SDPO: Segment-Level Direct Preference Optimization for Social Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kong, Aobo, Ma, Wentao, Zhao, Shiwan, Li, Yongbin, Wu, Yuchuan, Wang, Ke, Liu, Xiaoqian, Li, Qicheng, Qin, Yong, Huang, Fei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917938436505600
author Kong, Aobo
Ma, Wentao
Zhao, Shiwan
Li, Yongbin
Wu, Yuchuan
Wang, Ke
Liu, Xiaoqian
Li, Qicheng
Qin, Yong
Huang, Fei
author_facet Kong, Aobo
Ma, Wentao
Zhao, Shiwan
Li, Yongbin
Wu, Yuchuan
Wang, Ke
Liu, Xiaoqian
Li, Qicheng
Qin, Yong
Huang, Fei
contents Social agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling complex social dialogues. Direct Preference Optimization (DPO) has proven effective in aligning LLM behavior with human preferences across various agent tasks. However, standard DPO focuses solely on individual turns, which limits its effectiveness in multi-turn social interactions. Several DPO-based multi-turn alignment methods with session-level data have shown potential in addressing this problem.While these methods consider multiple turns across entire sessions, they are often overly coarse-grained, introducing training noise, and lack robust theoretical support. To resolve these limitations, we propose Segment-Level Direct Preference Optimization (SDPO), which dynamically select key segments within interactions to optimize multi-turn agent behavior. SDPO minimizes training noise and is grounded in a rigorous theoretical framework. Evaluations on the SOTOPIA benchmark demonstrate that SDPO-tuned agents consistently outperform both existing DPO-based methods and proprietary LLMs like GPT-4o, underscoring SDPO's potential to advance the social intelligence of LLM-based agents. We release our code and data at https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/SDPO.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SDPO: Segment-Level Direct Preference Optimization for Social Agents
Kong, Aobo
Ma, Wentao
Zhao, Shiwan
Li, Yongbin
Wu, Yuchuan
Wang, Ke
Liu, Xiaoqian
Li, Qicheng
Qin, Yong
Huang, Fei
Artificial Intelligence
Computation and Language
Social agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling complex social dialogues. Direct Preference Optimization (DPO) has proven effective in aligning LLM behavior with human preferences across various agent tasks. However, standard DPO focuses solely on individual turns, which limits its effectiveness in multi-turn social interactions. Several DPO-based multi-turn alignment methods with session-level data have shown potential in addressing this problem.While these methods consider multiple turns across entire sessions, they are often overly coarse-grained, introducing training noise, and lack robust theoretical support. To resolve these limitations, we propose Segment-Level Direct Preference Optimization (SDPO), which dynamically select key segments within interactions to optimize multi-turn agent behavior. SDPO minimizes training noise and is grounded in a rigorous theoretical framework. Evaluations on the SOTOPIA benchmark demonstrate that SDPO-tuned agents consistently outperform both existing DPO-based methods and proprietary LLMs like GPT-4o, underscoring SDPO's potential to advance the social intelligence of LLM-based agents. We release our code and data at https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/SDPO.
title SDPO: Segment-Level Direct Preference Optimization for Social Agents
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2501.01821