HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yunwen, Hong, Shiyong, Xiao, Xijun, Jin, Jinqiu, Luo, Xuanyuan, Wang, Zhe, Chai, Zheng, Wu, Shikang, Zheng, Yuchao, Lin, Jingjian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914274548383744
author Huang, Yunwen
Hong, Shiyong
Xiao, Xijun
Jin, Jinqiu
Luo, Xuanyuan
Wang, Zhe
Chai, Zheng
Wu, Shikang
Zheng, Yuchao
Lin, Jingjian
author_facet Huang, Yunwen
Hong, Shiyong
Xiao, Xijun
Jin, Jinqiu
Luo, Xuanyuan
Wang, Zhe
Chai, Zheng
Wu, Shikang
Zheng, Yuchao
Lin, Jingjian
contents Industrial large-scale recommendation models (LRMs) face the challenge of jointly modeling long-range user behavior sequences and heterogeneous non-sequential features under strict efficiency constraints. However, most existing architectures employ a decoupled pipeline: long sequences are first compressed with a query-token based sequence compressor like LONGER, followed by fusion with dense features through token-mixing modules like RankMixer, which thereby limits both the representation capacity and the interaction flexibility. This paper presents HyFormer, a unified hybrid transformer architecture that tightly integrates long-sequence modeling and feature interaction into a single backbone. From the perspective of sequence modeling, we revisit and redesign query tokens in LRMs, and frame the LRM modeling task as an alternating optimization process that integrates two core components: Query Decoding which expands non-sequential features into Global Tokens and performs long sequence decoding over layer-wise key-value representations of long behavioral sequences; and Query Boosting which enhances cross-query and cross-sequence heterogeneous interactions via efficient token mixing. The two complementary mechanisms are performed iteratively to refine semantic representations across layers. Extensive experiments on billion-scale industrial datasets demonstrate that HyFormer consistently outperforms strong LONGER and RankMixer baselines under comparable parameter and FLOPs budgets, while exhibiting superior scaling behavior with increasing parameters and FLOPs. Large-scale online A/B tests in high-traffic production systems further validate its effectiveness, showing significant gains over deployed state-of-the-art models. These results highlight the practicality and scalability of HyFormer as a unified modeling framework for industrial LRMs.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12681
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction
Huang, Yunwen
Hong, Shiyong
Xiao, Xijun
Jin, Jinqiu
Luo, Xuanyuan
Wang, Zhe
Chai, Zheng
Wu, Shikang
Zheng, Yuchao
Lin, Jingjian
Information Retrieval
Industrial large-scale recommendation models (LRMs) face the challenge of jointly modeling long-range user behavior sequences and heterogeneous non-sequential features under strict efficiency constraints. However, most existing architectures employ a decoupled pipeline: long sequences are first compressed with a query-token based sequence compressor like LONGER, followed by fusion with dense features through token-mixing modules like RankMixer, which thereby limits both the representation capacity and the interaction flexibility. This paper presents HyFormer, a unified hybrid transformer architecture that tightly integrates long-sequence modeling and feature interaction into a single backbone. From the perspective of sequence modeling, we revisit and redesign query tokens in LRMs, and frame the LRM modeling task as an alternating optimization process that integrates two core components: Query Decoding which expands non-sequential features into Global Tokens and performs long sequence decoding over layer-wise key-value representations of long behavioral sequences; and Query Boosting which enhances cross-query and cross-sequence heterogeneous interactions via efficient token mixing. The two complementary mechanisms are performed iteratively to refine semantic representations across layers. Extensive experiments on billion-scale industrial datasets demonstrate that HyFormer consistently outperforms strong LONGER and RankMixer baselines under comparable parameter and FLOPs budgets, while exhibiting superior scaling behavior with increasing parameters and FLOPs. Large-scale online A/B tests in high-traffic production systems further validate its effectiveness, showing significant gains over deployed state-of-the-art models. These results highlight the practicality and scalability of HyFormer as a unified modeling framework for industrial LRMs.
title HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction
topic Information Retrieval
url https://arxiv.org/abs/2601.12681