Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Dayu, Monaikul, Natawut, Ding, Amanda, Tan, Bozhao, Mosaliganti, Kishore, Iyengar, Giri
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2411.03356
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915007034294272
author Yang, Dayu
Monaikul, Natawut
Ding, Amanda
Tan, Bozhao
Mosaliganti, Kishore
Iyengar, Giri
author_facet Yang, Dayu
Monaikul, Natawut
Ding, Amanda
Tan, Bozhao
Mosaliganti, Kishore
Iyengar, Giri
contents In the era of data-driven decision-making, accurate table-level representations and efficient table recommendation systems are becoming increasingly crucial for improving table management, discovery, and analysis. However, existing approaches to tabular data representation often face limitations, primarily due to their focus on cell-level tasks and the lack of high-quality training data. To address these challenges, we first formulate a clear definition of table similarity in the context of data transformation activities within data-driven enterprises. This definition serves as the foundation for synthetic data generation, which require a well-defined data generation process. Building on this, we propose a novel synthetic data generation pipeline that harnesses the code generation and data manipulation capabilities of Large Language Models (LLMs) to create a large-scale synthetic dataset tailored for table-level representation learning. Through manual validation and performance comparisons on the table recommendation task, we demonstrate that the synthetic data generated by our pipeline aligns with our proposed definition of table similarity and significantly enhances table representations, leading to improved recommendation performance.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03356
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Table Representations with LLM-powered Synthetic Data Generation
Yang, Dayu
Monaikul, Natawut
Ding, Amanda
Tan, Bozhao
Mosaliganti, Kishore
Iyengar, Giri
Machine Learning
Artificial Intelligence
In the era of data-driven decision-making, accurate table-level representations and efficient table recommendation systems are becoming increasingly crucial for improving table management, discovery, and analysis. However, existing approaches to tabular data representation often face limitations, primarily due to their focus on cell-level tasks and the lack of high-quality training data. To address these challenges, we first formulate a clear definition of table similarity in the context of data transformation activities within data-driven enterprises. This definition serves as the foundation for synthetic data generation, which require a well-defined data generation process. Building on this, we propose a novel synthetic data generation pipeline that harnesses the code generation and data manipulation capabilities of Large Language Models (LLMs) to create a large-scale synthetic dataset tailored for table-level representation learning. Through manual validation and performance comparisons on the table recommendation task, we demonstrate that the synthetic data generated by our pipeline aligns with our proposed definition of table similarity and significantly enhances table representations, leading to improved recommendation performance.
title Enhancing Table Representations with LLM-powered Synthetic Data Generation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2411.03356