Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Maximillian, Sun, Ruoxi, Pfister, Tomas, Arık, Sercan Ö.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916865524105216
author Chen, Maximillian
Sun, Ruoxi
Pfister, Tomas
Arık, Sercan Ö.
author_facet Chen, Maximillian
Sun, Ruoxi
Pfister, Tomas
Arık, Sercan Ö.
contents Large language models (LLMs), optimized through human feedback, have rapidly emerged as a leading paradigm for developing intelligent conversational assistants. However, despite their strong performance across many benchmarks, LLM-based agents might still lack conversational skills such as disambiguation -- when they are faced with ambiguity, they often overhedge or implicitly guess users' true intents rather than asking clarification questions. Under task-specific settings, high-quality conversation samples are often limited, constituting a bottleneck for LLMs' ability to learn optimal dialogue action policies. We propose Action-Based Contrastive Self-Training (ACT), a quasi-online preference optimization algorithm based on Direct Preference Optimization (DPO), that enables data-efficient dialogue policy learning in multi-turn conversation modeling. We demonstrate ACT's efficacy under in data-efficient tuning scenarios, even when there is no action label available, using multiple real-world conversational tasks: tabular-grounded question-answering, machine reading comprehension, and AmbigSQL, a novel task for disambiguating information-seeking requests for complex SQL generation towards data analysis agents. Additionally, we propose evaluating LLMs' ability to function as conversational agents by examining whether they can implicitly recognize and reason about ambiguity in conversation. ACT demonstrates substantial conversation modeling improvements over standard tuning approaches like supervised fine-tuning and DPO.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00222
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
Chen, Maximillian
Sun, Ruoxi
Pfister, Tomas
Arık, Sercan Ö.
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs), optimized through human feedback, have rapidly emerged as a leading paradigm for developing intelligent conversational assistants. However, despite their strong performance across many benchmarks, LLM-based agents might still lack conversational skills such as disambiguation -- when they are faced with ambiguity, they often overhedge or implicitly guess users' true intents rather than asking clarification questions. Under task-specific settings, high-quality conversation samples are often limited, constituting a bottleneck for LLMs' ability to learn optimal dialogue action policies. We propose Action-Based Contrastive Self-Training (ACT), a quasi-online preference optimization algorithm based on Direct Preference Optimization (DPO), that enables data-efficient dialogue policy learning in multi-turn conversation modeling. We demonstrate ACT's efficacy under in data-efficient tuning scenarios, even when there is no action label available, using multiple real-world conversational tasks: tabular-grounded question-answering, machine reading comprehension, and AmbigSQL, a novel task for disambiguating information-seeking requests for complex SQL generation towards data analysis agents. Additionally, we propose evaluating LLMs' ability to function as conversational agents by examining whether they can implicitly recognize and reason about ambiguity in conversation. ACT demonstrates substantial conversation modeling improvements over standard tuning approaches like supervised fine-tuning and DPO.
title Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.00222