From Scaling to Structured Expressivity: Rethinking Transformers for CTR Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Bencheng, Lei, Yuejie, Zeng, Zhiyuan, Deng, Zheye, Wang, Di, Lin, Kaiyi, Wang, Pengjie, Yu, Chuan, Xu, Jian, Zheng, Bo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914621100654592
author Yan, Bencheng
Lei, Yuejie
Zeng, Zhiyuan
Deng, Zheye
Wang, Di
Lin, Kaiyi
Wang, Pengjie
Yu, Chuan
Xu, Jian
Zheng, Bo
author_facet Yan, Bencheng
Lei, Yuejie
Zeng, Zhiyuan
Deng, Zheye
Wang, Di
Lin, Kaiyi
Wang, Pengjie
Yu, Chuan
Xu, Jian
Zheng, Bo
contents Despite massive investments in scale, deep models for click-through rate (CTR) prediction often exhibit rapidly diminishing returns -- a stark contrast to the {predictable scaling laws} seen in large language models (LLMs). We identify the root cause as a {fundamental} \textit{structural misalignment}: {standard} Transformers assume sequential compositionality, whereas CTR data demand combinatorial reasoning over {heterogeneous} fields. To restore alignment, we introduce the \textbf{Field-Aware Transformer (FAT)}. {By reconstructing the standard Transformer block with field-centric parameters, FAT achieves \textit{structured expressivity}, {fundamentally shifting the model complexity dependence from the total vocabulary size $n$ with the number of fields $F$ ($n \gg F$).}} Crucially, to decouple model capacity from field cardinality, FAT employs a {Basis-Composed Hypernetwork} to synthesize field-specific parameters from shared bases, further reducing parameter complexity. {Theoretically, we ground this scaling behavior through a formal scaling law based on Rademacher complexity. Empirically, FAT outperforms exisiting state-of-the-art methods with up to \textbf{+4.38\%} AUC improvement, and delivers \textbf{+2.33\%} CTR and \textbf{+0.66\%} RPM in live production.} Our work establishes that scalable recommendation arises not from size alone, but from \textit{structured expressivity} -- architectural coherence with data semantics.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12081
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Scaling to Structured Expressivity: Rethinking Transformers for CTR Prediction
Yan, Bencheng
Lei, Yuejie
Zeng, Zhiyuan
Deng, Zheye
Wang, Di
Lin, Kaiyi
Wang, Pengjie
Yu, Chuan
Xu, Jian
Zheng, Bo
Information Retrieval
Machine Learning
Despite massive investments in scale, deep models for click-through rate (CTR) prediction often exhibit rapidly diminishing returns -- a stark contrast to the {predictable scaling laws} seen in large language models (LLMs). We identify the root cause as a {fundamental} \textit{structural misalignment}: {standard} Transformers assume sequential compositionality, whereas CTR data demand combinatorial reasoning over {heterogeneous} fields. To restore alignment, we introduce the \textbf{Field-Aware Transformer (FAT)}. {By reconstructing the standard Transformer block with field-centric parameters, FAT achieves \textit{structured expressivity}, {fundamentally shifting the model complexity dependence from the total vocabulary size $n$ with the number of fields $F$ ($n \gg F$).}} Crucially, to decouple model capacity from field cardinality, FAT employs a {Basis-Composed Hypernetwork} to synthesize field-specific parameters from shared bases, further reducing parameter complexity. {Theoretically, we ground this scaling behavior through a formal scaling law based on Rademacher complexity. Empirically, FAT outperforms exisiting state-of-the-art methods with up to \textbf{+4.38\%} AUC improvement, and delivers \textbf{+2.33\%} CTR and \textbf{+0.66\%} RPM in live production.} Our work establishes that scalable recommendation arises not from size alone, but from \textit{structured expressivity} -- architectural coherence with data semantics.
title From Scaling to Structured Expressivity: Rethinking Transformers for CTR Prediction
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2511.12081