Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Klenitskiy, Anton, Polev, Konstantin, Denisova, Daria, Vasilev, Alexey, Simakov, Dmitry, Gusev, Gleb
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917277274734592
author Klenitskiy, Anton
Polev, Konstantin
Denisova, Daria
Vasilev, Alexey
Simakov, Dmitry
Gusev, Gleb
author_facet Klenitskiy, Anton
Polev, Konstantin
Denisova, Daria
Vasilev, Alexey
Simakov, Dmitry
Gusev, Gleb
contents Many current state-of-the-art models for sequential recommendations are based on transformer architectures. Interpretation and explanation of such black box models is an important research question, as a better understanding of their internals can help understand, influence, and control their behavior, which is very important in a variety of real-world applications. Recently, sparse autoencoders (SAE) have been shown to be a promising unsupervised approach to extract interpretable features from neural networks. In this work, we extend SAE to sequential recommender systems and propose a framework for interpreting and controlling model representations. We show that this approach can be successfully applied to the transformer trained on a sequential recommendation task: directions learned in such an unsupervised regime turn out to be more interpretable and monosemantic than the original hidden state dimensions. Further, we demonstrate a straightforward way to effectively and flexibly control the model's behavior, giving developers and users of recommendation systems the ability to adjust their recommendations to various custom scenarios and contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12202
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control
Klenitskiy, Anton
Polev, Konstantin
Denisova, Daria
Vasilev, Alexey
Simakov, Dmitry
Gusev, Gleb
Information Retrieval
Artificial Intelligence
Machine Learning
Many current state-of-the-art models for sequential recommendations are based on transformer architectures. Interpretation and explanation of such black box models is an important research question, as a better understanding of their internals can help understand, influence, and control their behavior, which is very important in a variety of real-world applications. Recently, sparse autoencoders (SAE) have been shown to be a promising unsupervised approach to extract interpretable features from neural networks. In this work, we extend SAE to sequential recommender systems and propose a framework for interpreting and controlling model representations. We show that this approach can be successfully applied to the transformer trained on a sequential recommendation task: directions learned in such an unsupervised regime turn out to be more interpretable and monosemantic than the original hidden state dimensions. Further, we demonstrate a straightforward way to effectively and flexibly control the model's behavior, giving developers and users of recommendation systems the ability to adjust their recommendations to various custom scenarios and contexts.
title Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control
topic Information Retrieval
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.12202