ViMQ: A Vietnamese Medical Question Dataset for Healthcare Dialogue System Development

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huy, Ta Duc, Tu, Nguyen Anh, Vu, Tran Hoang, Minh, Nguyen Phuc, Phan, Nguyen, Bui, Trung H., Truong, Steven Q. H.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910411466473472
author Huy, Ta Duc
Tu, Nguyen Anh
Vu, Tran Hoang
Minh, Nguyen Phuc
Phan, Nguyen
Bui, Trung H.
Truong, Steven Q. H.
author_facet Huy, Ta Duc
Tu, Nguyen Anh
Vu, Tran Hoang
Minh, Nguyen Phuc
Phan, Nguyen
Bui, Trung H.
Truong, Steven Q. H.
contents Existing medical text datasets usually take the form of question and answer pairs that support the task of natural language generation, but lacking the composite annotations of the medical terms. In this study, we publish a Vietnamese dataset of medical questions from patients with sentence-level and entity-level annotations for the Intent Classification and Named Entity Recognition tasks. The tag sets for two tasks are in medical domain and can facilitate the development of task-oriented healthcare chatbots with better comprehension of queries from patients. We train baseline models for the two tasks and propose a simple self-supervised training strategy with span-noise modelling that substantially improves the performance. Dataset and code will be published at https://github.com/tadeephuy/ViMQ
format Preprint
id arxiv_https___arxiv_org_abs_2304_14405
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ViMQ: A Vietnamese Medical Question Dataset for Healthcare Dialogue System Development
Huy, Ta Duc
Tu, Nguyen Anh
Vu, Tran Hoang
Minh, Nguyen Phuc
Phan, Nguyen
Bui, Trung H.
Truong, Steven Q. H.
Computation and Language
Existing medical text datasets usually take the form of question and answer pairs that support the task of natural language generation, but lacking the composite annotations of the medical terms. In this study, we publish a Vietnamese dataset of medical questions from patients with sentence-level and entity-level annotations for the Intent Classification and Named Entity Recognition tasks. The tag sets for two tasks are in medical domain and can facilitate the development of task-oriented healthcare chatbots with better comprehension of queries from patients. We train baseline models for the two tasks and propose a simple self-supervised training strategy with span-noise modelling that substantially improves the performance. Dataset and code will be published at https://github.com/tadeephuy/ViMQ
title ViMQ: A Vietnamese Medical Question Dataset for Healthcare Dialogue System Development
topic Computation and Language
url https://arxiv.org/abs/2304.14405