LLM Embeddings for Deep Learning on Tabular Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Koloski, Boshko, Margeloiu, Andrei, Jiang, Xiangjian, Škrlj, Blaž, Simidjievski, Nikola, Jamnik, Mateja
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910831573204992
author Koloski, Boshko
Margeloiu, Andrei
Jiang, Xiangjian
Škrlj, Blaž
Simidjievski, Nikola
Jamnik, Mateja
author_facet Koloski, Boshko
Margeloiu, Andrei
Jiang, Xiangjian
Škrlj, Blaž
Simidjievski, Nikola
Jamnik, Mateja
contents Tabular deep-learning methods require embedding numerical and categorical input features into high-dimensional spaces before processing them. Existing methods deal with this heterogeneous nature of tabular data by employing separate type-specific encoding approaches. This limits the cross-table transfer potential and the exploitation of pre-trained knowledge. We propose a novel approach that first transforms tabular data into text, and then leverages pre-trained representations from LLMs to encode this data, resulting in a plug-and-play solution to improv ing deep-learning tabular methods. We demonstrate that our approach improves accuracy over competitive models, such as MLP, ResNet and FT-Transformer, by validating on seven classification datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11596
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Embeddings for Deep Learning on Tabular Data
Koloski, Boshko
Margeloiu, Andrei
Jiang, Xiangjian
Škrlj, Blaž
Simidjievski, Nikola
Jamnik, Mateja
Machine Learning
Artificial Intelligence
Tabular deep-learning methods require embedding numerical and categorical input features into high-dimensional spaces before processing them. Existing methods deal with this heterogeneous nature of tabular data by employing separate type-specific encoding approaches. This limits the cross-table transfer potential and the exploitation of pre-trained knowledge. We propose a novel approach that first transforms tabular data into text, and then leverages pre-trained representations from LLMs to encode this data, resulting in a plug-and-play solution to improv ing deep-learning tabular methods. We demonstrate that our approach improves accuracy over competitive models, such as MLP, ResNet and FT-Transformer, by validating on seven classification datasets.
title LLM Embeddings for Deep Learning on Tabular Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.11596