TREB: a BERT attempt for imputing tabular data imputation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shuyue, Zhou, Wenjun, Jiang, Han drk-m-s, Wang, Shuo, Zheng, Ren
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909331232915456
author Wang, Shuyue
Zhou, Wenjun
Jiang, Han drk-m-s
Wang, Shuo
Zheng, Ren
author_facet Wang, Shuyue
Zhou, Wenjun
Jiang, Han drk-m-s
Wang, Shuo
Zheng, Ren
contents TREB, a novel tabular imputation framework utilizing BERT, introduces a groundbreaking approach for handling missing values in tabular data. Unlike traditional methods that often overlook the specific demands of imputation, TREB leverages the robust capabilities of BERT to address this critical task. While many BERT-based approaches for tabular data have emerged, they frequently under-utilize the language model's full potential. To rectify this, TREB employs a BERT-based model fine-tuned specifically for the task of imputing real-valued continuous numbers in tabular datasets. The paper comprehensively addresses the unique challenges posed by tabular data imputation, emphasizing the importance of context-based interconnections. The effectiveness of TREB is validated through rigorous evaluation using the California Housing dataset. The results demonstrate its ability to preserve feature interrelationships and accurately impute missing values. Moreover, the authors shed light on the computational efficiency and environmental impact of TREB, quantifying the floating-point operations (FLOPs) and carbon footprint associated with its training and deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00022
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TREB: a BERT attempt for imputing tabular data imputation
Wang, Shuyue
Zhou, Wenjun
Jiang, Han drk-m-s
Wang, Shuo
Zheng, Ren
Machine Learning
TREB, a novel tabular imputation framework utilizing BERT, introduces a groundbreaking approach for handling missing values in tabular data. Unlike traditional methods that often overlook the specific demands of imputation, TREB leverages the robust capabilities of BERT to address this critical task. While many BERT-based approaches for tabular data have emerged, they frequently under-utilize the language model's full potential. To rectify this, TREB employs a BERT-based model fine-tuned specifically for the task of imputing real-valued continuous numbers in tabular datasets. The paper comprehensively addresses the unique challenges posed by tabular data imputation, emphasizing the importance of context-based interconnections. The effectiveness of TREB is validated through rigorous evaluation using the California Housing dataset. The results demonstrate its ability to preserve feature interrelationships and accurately impute missing values. Moreover, the authors shed light on the computational efficiency and environmental impact of TREB, quantifying the floating-point operations (FLOPs) and carbon footprint associated with its training and deployment.
title TREB: a BERT attempt for imputing tabular data imputation
topic Machine Learning
url https://arxiv.org/abs/2410.00022