Fostering Natural Conversation in Large Language Models with NICO: a Natural Interactive COnversation dataset

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sun, Renliang, Liu, Mengyuan, Yang, Shiping, Wang, Rui, He, Junqing, Zhang, Jiaxing
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914972440723456
author Sun, Renliang
Liu, Mengyuan
Yang, Shiping
Wang, Rui
He, Junqing
Zhang, Jiaxing
author_facet Sun, Renliang
Liu, Mengyuan
Yang, Shiping
Wang, Rui
He, Junqing
Zhang, Jiaxing
contents Benefiting from diverse instruction datasets, contemporary Large Language Models (LLMs) perform effectively as AI assistants in collaborating with humans. However, LLMs still struggle to generate natural and colloquial responses in real-world applications such as chatbots and psychological counseling that require more human-like interactions. To address these limitations, we introduce NICO, a Natural Interactive COnversation dataset in Chinese. We first use GPT-4-turbo to generate dialogue drafts and make them cover 20 daily-life topics and 5 types of social interactions. Then, we hire workers to revise these dialogues to ensure that they are free of grammatical errors and unnatural utterances. We define two dialogue-level natural conversation tasks and two sentence-level tasks for identifying and rewriting unnatural sentences. Multiple open-source and closed-source LLMs are tested and analyzed in detail. The experimental results highlight the challenge of the tasks and demonstrate how NICO can help foster the natural dialogue capabilities of LLMs. The dataset will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2408_09330
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fostering Natural Conversation in Large Language Models with NICO: a Natural Interactive COnversation dataset
Sun, Renliang
Liu, Mengyuan
Yang, Shiping
Wang, Rui
He, Junqing
Zhang, Jiaxing
Computation and Language
Benefiting from diverse instruction datasets, contemporary Large Language Models (LLMs) perform effectively as AI assistants in collaborating with humans. However, LLMs still struggle to generate natural and colloquial responses in real-world applications such as chatbots and psychological counseling that require more human-like interactions. To address these limitations, we introduce NICO, a Natural Interactive COnversation dataset in Chinese. We first use GPT-4-turbo to generate dialogue drafts and make them cover 20 daily-life topics and 5 types of social interactions. Then, we hire workers to revise these dialogues to ensure that they are free of grammatical errors and unnatural utterances. We define two dialogue-level natural conversation tasks and two sentence-level tasks for identifying and rewriting unnatural sentences. Multiple open-source and closed-source LLMs are tested and analyzed in detail. The experimental results highlight the challenge of the tasks and demonstrate how NICO can help foster the natural dialogue capabilities of LLMs. The dataset will be released.
title Fostering Natural Conversation in Large Language Models with NICO: a Natural Interactive COnversation dataset
topic Computation and Language
url https://arxiv.org/abs/2408.09330