Unifying Structured Data as Graph for Data-to-Text Pre-Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Shujie, Li, Liang, Geng, Ruiying, Yang, Min, Li, Binhua, Yuan, Guanghu, He, Wanwei, Yuan, Shao, Ma, Can, Huang, Fei, Li, Yongbin
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916079623733248
author Li, Shujie
Li, Liang
Geng, Ruiying
Yang, Min
Li, Binhua
Yuan, Guanghu
He, Wanwei
Yuan, Shao
Ma, Can
Huang, Fei
Li, Yongbin
author_facet Li, Shujie
Li, Liang
Geng, Ruiying
Yang, Min
Li, Binhua
Yuan, Guanghu
He, Wanwei
Yuan, Shao
Ma, Can
Huang, Fei
Li, Yongbin
contents Data-to-text (D2T) generation aims to transform structured data into natural language text. Data-to-text pre-training has proved to be powerful in enhancing D2T generation and yields impressive performances. However, previous pre-training methods either oversimplified structured data into a sequence without considering input structures or designed training objectives tailored for a specific data structure (e.g., table or knowledge graph). In this paper, we unify different types of structured data (i.e., table, key-value data, knowledge graph) into the graph format and cast different data-to-text generation tasks as graph-to-text generation. To effectively exploit the structural information of the input graph, we propose a structure-enhanced pre-training method for D2T generation by designing a structure-enhanced Transformer. Concretely, we devise a position matrix for the Transformer, encoding relative positional information of connected nodes in the input graph. In addition, we propose a new attention matrix to incorporate graph structures into the original Transformer by taking the available explicit connectivity structure into account. Extensive experiments on six benchmark datasets show the effectiveness of our model. Our source codes are available at https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/unid2t.
format Preprint
id arxiv_https___arxiv_org_abs_2401_01183
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unifying Structured Data as Graph for Data-to-Text Pre-Training
Li, Shujie
Li, Liang
Geng, Ruiying
Yang, Min
Li, Binhua
Yuan, Guanghu
He, Wanwei
Yuan, Shao
Ma, Can
Huang, Fei
Li, Yongbin
Computation and Language
Artificial Intelligence
Data-to-text (D2T) generation aims to transform structured data into natural language text. Data-to-text pre-training has proved to be powerful in enhancing D2T generation and yields impressive performances. However, previous pre-training methods either oversimplified structured data into a sequence without considering input structures or designed training objectives tailored for a specific data structure (e.g., table or knowledge graph). In this paper, we unify different types of structured data (i.e., table, key-value data, knowledge graph) into the graph format and cast different data-to-text generation tasks as graph-to-text generation. To effectively exploit the structural information of the input graph, we propose a structure-enhanced pre-training method for D2T generation by designing a structure-enhanced Transformer. Concretely, we devise a position matrix for the Transformer, encoding relative positional information of connected nodes in the input graph. In addition, we propose a new attention matrix to incorporate graph structures into the original Transformer by taking the available explicit connectivity structure into account. Extensive experiments on six benchmark datasets show the effectiveness of our model. Our source codes are available at https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/unid2t.
title Unifying Structured Data as Graph for Data-to-Text Pre-Training
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.01183