LLM4ES: Learning User Embeddings from Event Sequences via Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shestov, Aleksei, Zoloev, Omar, Makarenko, Maksim, Orlov, Mikhail, Fadeev, Egor, Kireev, Ivan, Savchenko, Andrey
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918251489918976
author Shestov, Aleksei
Zoloev, Omar
Makarenko, Maksim
Orlov, Mikhail
Fadeev, Egor
Kireev, Ivan
Savchenko, Andrey
author_facet Shestov, Aleksei
Zoloev, Omar
Makarenko, Maksim
Orlov, Mikhail
Fadeev, Egor
Kireev, Ivan
Savchenko, Andrey
contents This paper presents LLM4ES, a novel framework that exploits large pre-trained language models (LLMs) to derive user embeddings from event sequences. Event sequences are transformed into a textual representation, which is subsequently used to fine-tune an LLM through next-token prediction to generate high-quality embeddings. We introduce a text enrichment technique that enhances LLM adaptation to event sequence data, improving representation quality for low-variability domains. Experimental results demonstrate that LLM4ES achieves state-of-the-art performance in user classification tasks in financial and other domains, outperforming existing embedding methods. The resulting user embeddings can be incorporated into a wide range of applications, from user segmentation in finance to patient outcome prediction in healthcare.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05688
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM4ES: Learning User Embeddings from Event Sequences via Large Language Models
Shestov, Aleksei
Zoloev, Omar
Makarenko, Maksim
Orlov, Mikhail
Fadeev, Egor
Kireev, Ivan
Savchenko, Andrey
Information Retrieval
This paper presents LLM4ES, a novel framework that exploits large pre-trained language models (LLMs) to derive user embeddings from event sequences. Event sequences are transformed into a textual representation, which is subsequently used to fine-tune an LLM through next-token prediction to generate high-quality embeddings. We introduce a text enrichment technique that enhances LLM adaptation to event sequence data, improving representation quality for low-variability domains. Experimental results demonstrate that LLM4ES achieves state-of-the-art performance in user classification tasks in financial and other domains, outperforming existing embedding methods. The resulting user embeddings can be incorporated into a wide range of applications, from user segmentation in finance to patient outcome prediction in healthcare.
title LLM4ES: Learning User Embeddings from Event Sequences via Large Language Models
topic Information Retrieval
url https://arxiv.org/abs/2508.05688