Factor Augmented Supervised Learning with Text Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Zhanye, Han, Yuefeng, Yu, Xiufan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912529666539520
author Luo, Zhanye
Han, Yuefeng
Yu, Xiufan
author_facet Luo, Zhanye
Han, Yuefeng
Yu, Xiufan
contents Large language models (LLMs) generate text embeddings from text data, producing vector representations that capture the semantic meaning and contextual relationships of words. However, the high dimensionality of these embeddings often impedes efficiency and drives up computational cost in downstream tasks. To address this, we propose AutoEncoder-Augmented Learning with Text (AEALT), a supervised, factor-augmented framework that incorporates dimension reduction directly into pre-trained LLM workflows. First, we extract embeddings from text documents; next, we pass them through a supervised augmented autoencoder to learn low-dimensional, task-relevant latent factors. By modeling the nonlinear structure of complex embeddings, AEALT outperforms conventional deep-learning approaches that rely on raw embeddings. We validate its broad applicability with extensive experiments on classification, anomaly detection, and prediction tasks using multiple real-world public datasets. Numerical results demonstrate that AEALT yields substantial gains over both vanilla embeddings and several standard dimension reduction methods.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06548
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Factor Augmented Supervised Learning with Text Embeddings
Luo, Zhanye
Han, Yuefeng
Yu, Xiufan
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) generate text embeddings from text data, producing vector representations that capture the semantic meaning and contextual relationships of words. However, the high dimensionality of these embeddings often impedes efficiency and drives up computational cost in downstream tasks. To address this, we propose AutoEncoder-Augmented Learning with Text (AEALT), a supervised, factor-augmented framework that incorporates dimension reduction directly into pre-trained LLM workflows. First, we extract embeddings from text documents; next, we pass them through a supervised augmented autoencoder to learn low-dimensional, task-relevant latent factors. By modeling the nonlinear structure of complex embeddings, AEALT outperforms conventional deep-learning approaches that rely on raw embeddings. We validate its broad applicability with extensive experiments on classification, anomaly detection, and prediction tasks using multiple real-world public datasets. Numerical results demonstrate that AEALT yields substantial gains over both vanilla embeddings and several standard dimension reduction methods.
title Factor Augmented Supervised Learning with Text Embeddings
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.06548