CharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Production

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nie, Yixin, Guan, Lin, Ma, Zhongyao, Gupta, Anchit, Zhou, Yipin, Li, Xiao, Zhou, Zhengping, Zeng, Raymond, Zhou, Gelin, Chu, Shigan, Thampi, Ajay, Mu, Wancen, Shuster, Nathan, Wang, Ketong, Chen, Lin, Brewer, Jason, Hu, Derek Hao, McCauley, Alexander, Weston, Jason, Park, Sem, Zhang, Na, Tang, Kevin
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918366692769792
author Nie, Yixin
Guan, Lin
Ma, Zhongyao
Gupta, Anchit
Zhou, Yipin
Li, Xiao
Zhou, Zhengping
Zeng, Raymond
Zhou, Gelin
Chu, Shigan
Thampi, Ajay
Mu, Wancen
Shuster, Nathan
Wang, Ketong
Chen, Lin
Brewer, Jason
Hu, Derek Hao
McCauley, Alexander
Weston, Jason
Park, Sem
Zhang, Na
Tang, Kevin
author_facet Nie, Yixin
Guan, Lin
Ma, Zhongyao
Gupta, Anchit
Zhou, Yipin
Li, Xiao
Zhou, Zhengping
Zeng, Raymond
Zhou, Gelin
Chu, Shigan
Thampi, Ajay
Mu, Wancen
Shuster, Nathan
Wang, Ketong
Chen, Lin
Brewer, Jason
Hu, Derek Hao
McCauley, Alexander
Weston, Jason
Park, Sem
Zhang, Na
Tang, Kevin
contents This report presents CharacterFlywheel, an iterative flywheel process for improving large language models (LLMs) in production social chat applications across Instagram, WhatsApp, and Messenger. Starting from LLaMA 3.1, we refined models across 15 generations using data from both internal and external real-user traffic. Through continuous deployments from July 2024 to April 2025, we conducted controlled 7-day A/B tests showing consistent engagement improvements: 7 of 8 newly deployed models demonstrated positive lift over the baseline, with the strongest performers achieving up to 8.8% improvement in engagement breadth and 19.4% in engagement depth. We also observed substantial gains in steerability, with instruction following increasing from 59.2% to 84.8% and instruction violations decreasing from 26.6% to 5.8%. We detail the CharacterFlywheel process which integrates data curation, reward modeling to estimate and interpolate the landscape of engagement metrics, supervised fine-tuning (SFT), reinforcement learning (RL), and both offline and online evaluation to ensure reliable progress at each optimization step. We also discuss our methods for overfitting prevention and navigating production dynamics at scale. These contributions advance the scientific rigor and understanding of LLMs in social applications serving millions of users.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01973
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Production
Nie, Yixin
Guan, Lin
Ma, Zhongyao
Gupta, Anchit
Zhou, Yipin
Li, Xiao
Zhou, Zhengping
Zeng, Raymond
Zhou, Gelin
Chu, Shigan
Thampi, Ajay
Mu, Wancen
Shuster, Nathan
Wang, Ketong
Chen, Lin
Brewer, Jason
Hu, Derek Hao
McCauley, Alexander
Weston, Jason
Park, Sem
Zhang, Na
Tang, Kevin
Computation and Language
Artificial Intelligence
Social and Information Networks
This report presents CharacterFlywheel, an iterative flywheel process for improving large language models (LLMs) in production social chat applications across Instagram, WhatsApp, and Messenger. Starting from LLaMA 3.1, we refined models across 15 generations using data from both internal and external real-user traffic. Through continuous deployments from July 2024 to April 2025, we conducted controlled 7-day A/B tests showing consistent engagement improvements: 7 of 8 newly deployed models demonstrated positive lift over the baseline, with the strongest performers achieving up to 8.8% improvement in engagement breadth and 19.4% in engagement depth. We also observed substantial gains in steerability, with instruction following increasing from 59.2% to 84.8% and instruction violations decreasing from 26.6% to 5.8%. We detail the CharacterFlywheel process which integrates data curation, reward modeling to estimate and interpolate the landscape of engagement metrics, supervised fine-tuning (SFT), reinforcement learning (RL), and both offline and online evaluation to ensure reliable progress at each optimization step. We also discuss our methods for overfitting prevention and navigating production dynamics at scale. These contributions advance the scientific rigor and understanding of LLMs in social applications serving millions of users.
title CharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Production
topic Computation and Language
Artificial Intelligence
Social and Information Networks
url https://arxiv.org/abs/2603.01973