Baichuan2-Sum: Instruction Finetune Baichuan2-7B Model for Dialogue Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Jianfei, Chen, Yancan, Ou, Yimin, Yu, Hanyi, Shu, Kai, Xiao, Yiyong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917629793402880
author Xiao, Jianfei
Chen, Yancan
Ou, Yimin
Yu, Hanyi
Shu, Kai
Xiao, Yiyong
author_facet Xiao, Jianfei
Chen, Yancan
Ou, Yimin
Yu, Hanyi
Shu, Kai
Xiao, Yiyong
contents Large language models (LLMs) like Llama, Baichuan and Bloom models show remarkable ability with instruction fine-tuning in many natural language tasks. Nevertheless, for the dialogue summarization task, which aims to generate summaries for different roles in dialogue, most of the state-of-the-art methods conduct on small models (e.g Bart and Bert). Existing methods try to add task specified optimization on small models like adding global-local centrality score to models. In this paper, we propose an instruction fine-tuning model: Baichuan2-Sum, for role-oriented diaglouge summarization. By setting different instructions for different roles, the model can learn from the dialogue interactions and output the expected summaries. Furthermore, we applied NEFTune technique to add suitable noise during training to improve the results. The experiments demonstrate that the proposed model achieves the new state-of-the-art results on two public dialogue summarization datasets: CSDS and SAMSUM. We release our model and related codes to facilitate future studies on dialogue summarization task.
format Preprint
id arxiv_https___arxiv_org_abs_2401_15496
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Baichuan2-Sum: Instruction Finetune Baichuan2-7B Model for Dialogue Summarization
Xiao, Jianfei
Chen, Yancan
Ou, Yimin
Yu, Hanyi
Shu, Kai
Xiao, Yiyong
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) like Llama, Baichuan and Bloom models show remarkable ability with instruction fine-tuning in many natural language tasks. Nevertheless, for the dialogue summarization task, which aims to generate summaries for different roles in dialogue, most of the state-of-the-art methods conduct on small models (e.g Bart and Bert). Existing methods try to add task specified optimization on small models like adding global-local centrality score to models. In this paper, we propose an instruction fine-tuning model: Baichuan2-Sum, for role-oriented diaglouge summarization. By setting different instructions for different roles, the model can learn from the dialogue interactions and output the expected summaries. Furthermore, we applied NEFTune technique to add suitable noise during training to improve the results. The experiments demonstrate that the proposed model achieves the new state-of-the-art results on two public dialogue summarization datasets: CSDS and SAMSUM. We release our model and related codes to facilitate future studies on dialogue summarization task.
title Baichuan2-Sum: Instruction Finetune Baichuan2-7B Model for Dialogue Summarization
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2401.15496