A Large-Scale Web Search Dataset for Federated Online Learning to Rank

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gregoriadis, Marcel, Kang, Jingwei, Pouwelse, Johan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915448735399936
author Gregoriadis, Marcel
Kang, Jingwei
Pouwelse, Johan
author_facet Gregoriadis, Marcel
Kang, Jingwei
Pouwelse, Johan
contents The centralized collection of search interaction logs for training ranking models raises significant privacy concerns. Federated Online Learning to Rank (FOLTR) offers a privacy-preserving alternative by enabling collaborative model training without sharing raw user data. However, benchmarks in FOLTR are largely based on random partitioning of classical learning-to-rank datasets, simulated user clicks, and the assumption of synchronous client participation. This oversimplifies real-world dynamics and undermines the realism of experimental results. We present AOL4FOLTR, a large-scale web search dataset with 2.6 million queries from 10,000 users. Our dataset addresses key limitations of existing benchmarks by including user identifiers, real click data, and query timestamps, enabling realistic user partitioning, behavior modeling, and asynchronous federated learning scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12353
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Large-Scale Web Search Dataset for Federated Online Learning to Rank
Gregoriadis, Marcel
Kang, Jingwei
Pouwelse, Johan
Information Retrieval
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
The centralized collection of search interaction logs for training ranking models raises significant privacy concerns. Federated Online Learning to Rank (FOLTR) offers a privacy-preserving alternative by enabling collaborative model training without sharing raw user data. However, benchmarks in FOLTR are largely based on random partitioning of classical learning-to-rank datasets, simulated user clicks, and the assumption of synchronous client participation. This oversimplifies real-world dynamics and undermines the realism of experimental results. We present AOL4FOLTR, a large-scale web search dataset with 2.6 million queries from 10,000 users. Our dataset addresses key limitations of existing benchmarks by including user identifiers, real click data, and query timestamps, enabling realistic user partitioning, behavior modeling, and asynchronous federated learning scenarios.
title A Large-Scale Web Search Dataset for Federated Online Learning to Rank
topic Information Retrieval
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2508.12353