LLM Embeddings for Deep Learning on Tabular Data
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910831573204992 |
|---|---|
| author | Koloski, Boshko Margeloiu, Andrei Jiang, Xiangjian Škrlj, Blaž Simidjievski, Nikola Jamnik, Mateja |
| author_facet | Koloski, Boshko Margeloiu, Andrei Jiang, Xiangjian Škrlj, Blaž Simidjievski, Nikola Jamnik, Mateja |
| contents | Tabular deep-learning methods require embedding numerical and categorical input features into high-dimensional spaces before processing them. Existing methods deal with this heterogeneous nature of tabular data by employing separate type-specific encoding approaches. This limits the cross-table transfer potential and the exploitation of pre-trained knowledge. We propose a novel approach that first transforms tabular data into text, and then leverages pre-trained representations from LLMs to encode this data, resulting in a plug-and-play solution to improv ing deep-learning tabular methods. We demonstrate that our approach improves accuracy over competitive models, such as MLP, ResNet and FT-Transformer, by validating on seven classification datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_11596 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LLM Embeddings for Deep Learning on Tabular Data Koloski, Boshko Margeloiu, Andrei Jiang, Xiangjian Škrlj, Blaž Simidjievski, Nikola Jamnik, Mateja Machine Learning Artificial Intelligence Tabular deep-learning methods require embedding numerical and categorical input features into high-dimensional spaces before processing them. Existing methods deal with this heterogeneous nature of tabular data by employing separate type-specific encoding approaches. This limits the cross-table transfer potential and the exploitation of pre-trained knowledge. We propose a novel approach that first transforms tabular data into text, and then leverages pre-trained representations from LLMs to encode this data, resulting in a plug-and-play solution to improv ing deep-learning tabular methods. We demonstrate that our approach improves accuracy over competitive models, such as MLP, ResNet and FT-Transformer, by validating on seven classification datasets. |
| title | LLM Embeddings for Deep Learning on Tabular Data |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2502.11596 |