VK-LSVD: A Large-Scale Industrial Dataset for Short-Video Recommendation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Poslavsky, Aleksandr, D'yakonov, Alexander, Dorn, Yuriy, Zimovnov, Andrey
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911436535496704
author Poslavsky, Aleksandr
D'yakonov, Alexander
Dorn, Yuriy
Zimovnov, Andrey
author_facet Poslavsky, Aleksandr
D'yakonov, Alexander
Dorn, Yuriy
Zimovnov, Andrey
contents Short-video recommendation presents unique challenges, such as modeling rapid user interest shifts from implicit feedback, but progress is constrained by a lack of large-scale open datasets that reflect real-world platform dynamics. To bridge this gap, we introduce the VK Large Short-Video Dataset (VK-LSVD), the largest publicly available industrial dataset of its kind. VK-LSVD offers an unprecedented scale of over 40 billion interactions from 10 million users and almost 20 million videos over six months, alongside rich features including content embeddings, diverse feedback signals, and contextual metadata. Our analysis supports the dataset's quality and diversity. The dataset's immediate impact is confirmed by its central role in the live VK RecSys Challenge 2025. VK-LSVD provides a vital, open dataset to use in building realistic benchmarks to accelerate research in sequential recommendation, cold-start scenarios, and next-generation recommender systems.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04567
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VK-LSVD: A Large-Scale Industrial Dataset for Short-Video Recommendation
Poslavsky, Aleksandr
D'yakonov, Alexander
Dorn, Yuriy
Zimovnov, Andrey
Information Retrieval
Computers and Society
H.3.3; H.3.5; I.2.6; H.2.8; K.4.1
Short-video recommendation presents unique challenges, such as modeling rapid user interest shifts from implicit feedback, but progress is constrained by a lack of large-scale open datasets that reflect real-world platform dynamics. To bridge this gap, we introduce the VK Large Short-Video Dataset (VK-LSVD), the largest publicly available industrial dataset of its kind. VK-LSVD offers an unprecedented scale of over 40 billion interactions from 10 million users and almost 20 million videos over six months, alongside rich features including content embeddings, diverse feedback signals, and contextual metadata. Our analysis supports the dataset's quality and diversity. The dataset's immediate impact is confirmed by its central role in the live VK RecSys Challenge 2025. VK-LSVD provides a vital, open dataset to use in building realistic benchmarks to accelerate research in sequential recommendation, cold-start scenarios, and next-generation recommender systems.
title VK-LSVD: A Large-Scale Industrial Dataset for Short-Video Recommendation
topic Information Retrieval
Computers and Society
H.3.3; H.3.5; I.2.6; H.2.8; K.4.1
url https://arxiv.org/abs/2602.04567