AdNanny: One Reasoning LLM for All Offline Ads Recommendation Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hu, Nan, Li, Han, Sun, Jimeng, Wang, Lu, Yang, Fangkai, Qiao, Bo, Zhao, Pu, Dai, David, Liu, Mengyu, Zhan, Yuefeng, Zhang, Jianjin, Han, Weihao, Sun, Allen, Lin, Qingwei, Rajmohan, Saravan, Zhang, Dongmei, Deng, Denvy, Sun, Feng, Zhang, Qi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915768393793536
author Hu, Nan
Li, Han
Sun, Jimeng
Wang, Lu
Yang, Fangkai
Qiao, Bo
Zhao, Pu
Dai, David
Liu, Mengyu
Zhan, Yuefeng
Zhang, Jianjin
Han, Weihao
Sun, Allen
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
Deng, Denvy
Sun, Feng
Zhang, Qi
author_facet Hu, Nan
Li, Han
Sun, Jimeng
Wang, Lu
Yang, Fangkai
Qiao, Bo
Zhao, Pu
Dai, David
Liu, Mengyu
Zhan, Yuefeng
Zhang, Jianjin
Han, Weihao
Sun, Allen
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
Deng, Denvy
Sun, Feng
Zhang, Qi
contents Large Language Models (LLMs) have shown strong capabilities in Natural Language Understanding and Generation, but deploying them directly in online advertising systems is often impractical due to strict millisecond-level latency constraints. This has motivated the use of LLMs offline to improve retrieval, ranking, and recommendation models. Existing solutions typically fine-tune separate LLMs for individual tasks such as query-ad relevance labeling, keyword-based query generation, and user profiling. This results in redundant models, high maintenance cost, and limited performance gains despite substantial overlap in domain knowledge and reasoning patterns. We introduce AdNanny, a unified reasoning-centric LLM that serves as a shared backbone for offline advertising tasks. AdNanny is obtained by fine-tuning a public 671B-parameter DeepSeek-R1 checkpoint using a scalable training system that supports hybrid dense-MoE parallelism. We construct reasoning-augmented corpora that pair structured supervision with step-by-step natural language explanations. A multi-task supervised fine-tuning stage with adaptive reweighting enables AdNanny to handle diverse labeling and generation tasks in a consistent reasoning format. This is followed by reinforcement learning using downstream advertising metrics to align model behavior with online retrieval and ranking objectives. AdNanny is deployed in production within Bing Ads, where it significantly reduces manual labeling effort and improves accuracy across multiple offline tasks. By consolidating many task-specific models into a single reasoning-centric foundation model, AdNanny provides a scalable and cost-effective solution for large-scale advertising systems.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01563
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AdNanny: One Reasoning LLM for All Offline Ads Recommendation Tasks
Hu, Nan
Li, Han
Sun, Jimeng
Wang, Lu
Yang, Fangkai
Qiao, Bo
Zhao, Pu
Dai, David
Liu, Mengyu
Zhan, Yuefeng
Zhang, Jianjin
Han, Weihao
Sun, Allen
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
Deng, Denvy
Sun, Feng
Zhang, Qi
Software Engineering
Large Language Models (LLMs) have shown strong capabilities in Natural Language Understanding and Generation, but deploying them directly in online advertising systems is often impractical due to strict millisecond-level latency constraints. This has motivated the use of LLMs offline to improve retrieval, ranking, and recommendation models. Existing solutions typically fine-tune separate LLMs for individual tasks such as query-ad relevance labeling, keyword-based query generation, and user profiling. This results in redundant models, high maintenance cost, and limited performance gains despite substantial overlap in domain knowledge and reasoning patterns. We introduce AdNanny, a unified reasoning-centric LLM that serves as a shared backbone for offline advertising tasks. AdNanny is obtained by fine-tuning a public 671B-parameter DeepSeek-R1 checkpoint using a scalable training system that supports hybrid dense-MoE parallelism. We construct reasoning-augmented corpora that pair structured supervision with step-by-step natural language explanations. A multi-task supervised fine-tuning stage with adaptive reweighting enables AdNanny to handle diverse labeling and generation tasks in a consistent reasoning format. This is followed by reinforcement learning using downstream advertising metrics to align model behavior with online retrieval and ranking objectives. AdNanny is deployed in production within Bing Ads, where it significantly reduces manual labeling effort and improves accuracy across multiple offline tasks. By consolidating many task-specific models into a single reasoning-centric foundation model, AdNanny provides a scalable and cost-effective solution for large-scale advertising systems.
title AdNanny: One Reasoning LLM for All Offline Ads Recommendation Tasks
topic Software Engineering
url https://arxiv.org/abs/2602.01563