Saved in:
Bibliographic Details
Main Authors: Cokbas, Mertcan, Liu, Ziteng, Tao, Zeyi, Veliz, Elder, Huang, Qin, Wen, Ellie, Li, Huayu, Jin, Qiang, Duman, Murat, Au, Benjamin, Lebanon, Guy, Chordia, Sagar, Zhang, Chengkai
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.02215
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913052617605120
author Cokbas, Mertcan
Liu, Ziteng
Tao, Zeyi
Veliz, Elder
Huang, Qin
Wen, Ellie
Li, Huayu
Jin, Qiang
Duman, Murat
Au, Benjamin
Lebanon, Guy
Chordia, Sagar
Zhang, Chengkai
author_facet Cokbas, Mertcan
Liu, Ziteng
Tao, Zeyi
Veliz, Elder
Huang, Qin
Wen, Ellie
Li, Huayu
Jin, Qiang
Duman, Murat
Au, Benjamin
Lebanon, Guy
Chordia, Sagar
Zhang, Chengkai
contents Training large-scale recommendation models under a single global objective implicitly assumes homogeneity across user populations. However, real-world data are composites of heterogeneous cohorts with distinct conditional distributions. As models increase in scale and complexity and as more data is used for training, they become dominated by central distribution patterns, neglecting head and tail regions. This imbalance limits the model's learning ability and can result in inactive attention weights or dead neurons. In this paper, we reveal how the attention mechanism can play a key role in factorization machines for shared embedding selection, and propose to address this challenge by analyzing the substructures in the dataset and exposing those with strong distributional contrast through auxiliary learning. Unlike previous research, which heuristically applies weighted labels or multi-task heads to mitigate such biases, we leverage partially conflicting auxiliary labels to regularize the shared representation. This approach customizes the learning process of attention layers to preserve mutual information with minority cohorts while improving global performance. We evaluated proposed method on massive production datasets with billions of data points each for six SOTA models. Experiments show that the factorization machine is able to capture fine-grained user-ad interactions using the proposed method, achieving up to a 0.16% reduction in normalized entropy overall and delivering gains exceeding 0.30% on targeted minority cohorts.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Large-Scale Recommender Systems with Auxiliary Learning
Cokbas, Mertcan
Liu, Ziteng
Tao, Zeyi
Veliz, Elder
Huang, Qin
Wen, Ellie
Li, Huayu
Jin, Qiang
Duman, Murat
Au, Benjamin
Lebanon, Guy
Chordia, Sagar
Zhang, Chengkai
Machine Learning
Training large-scale recommendation models under a single global objective implicitly assumes homogeneity across user populations. However, real-world data are composites of heterogeneous cohorts with distinct conditional distributions. As models increase in scale and complexity and as more data is used for training, they become dominated by central distribution patterns, neglecting head and tail regions. This imbalance limits the model's learning ability and can result in inactive attention weights or dead neurons. In this paper, we reveal how the attention mechanism can play a key role in factorization machines for shared embedding selection, and propose to address this challenge by analyzing the substructures in the dataset and exposing those with strong distributional contrast through auxiliary learning. Unlike previous research, which heuristically applies weighted labels or multi-task heads to mitigate such biases, we leverage partially conflicting auxiliary labels to regularize the shared representation. This approach customizes the learning process of attention layers to preserve mutual information with minority cohorts while improving global performance. We evaluated proposed method on massive production datasets with billions of data points each for six SOTA models. Experiments show that the factorization machine is able to capture fine-grained user-ad interactions using the proposed method, achieving up to a 0.16% reduction in normalized entropy overall and delivering gains exceeding 0.30% on targeted minority cohorts.
title Improving Large-Scale Recommender Systems with Auxiliary Learning
topic Machine Learning
url https://arxiv.org/abs/2510.02215