Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917277274734592 |
|---|---|
| author | Klenitskiy, Anton Polev, Konstantin Denisova, Daria Vasilev, Alexey Simakov, Dmitry Gusev, Gleb |
| author_facet | Klenitskiy, Anton Polev, Konstantin Denisova, Daria Vasilev, Alexey Simakov, Dmitry Gusev, Gleb |
| contents | Many current state-of-the-art models for sequential recommendations are based on transformer architectures. Interpretation and explanation of such black box models is an important research question, as a better understanding of their internals can help understand, influence, and control their behavior, which is very important in a variety of real-world applications. Recently, sparse autoencoders (SAE) have been shown to be a promising unsupervised approach to extract interpretable features from neural networks.
In this work, we extend SAE to sequential recommender systems and propose a framework for interpreting and controlling model representations. We show that this approach can be successfully applied to the transformer trained on a sequential recommendation task: directions learned in such an unsupervised regime turn out to be more interpretable and monosemantic than the original hidden state dimensions. Further, we demonstrate a straightforward way to effectively and flexibly control the model's behavior, giving developers and users of recommendation systems the ability to adjust their recommendations to various custom scenarios and contexts. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_12202 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control Klenitskiy, Anton Polev, Konstantin Denisova, Daria Vasilev, Alexey Simakov, Dmitry Gusev, Gleb Information Retrieval Artificial Intelligence Machine Learning Many current state-of-the-art models for sequential recommendations are based on transformer architectures. Interpretation and explanation of such black box models is an important research question, as a better understanding of their internals can help understand, influence, and control their behavior, which is very important in a variety of real-world applications. Recently, sparse autoencoders (SAE) have been shown to be a promising unsupervised approach to extract interpretable features from neural networks. In this work, we extend SAE to sequential recommender systems and propose a framework for interpreting and controlling model representations. We show that this approach can be successfully applied to the transformer trained on a sequential recommendation task: directions learned in such an unsupervised regime turn out to be more interpretable and monosemantic than the original hidden state dimensions. Further, we demonstrate a straightforward way to effectively and flexibly control the model's behavior, giving developers and users of recommendation systems the ability to adjust their recommendations to various custom scenarios and contexts. |
| title | Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control |
| topic | Information Retrieval Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2507.12202 |