Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Chengbing, Zhang, Yang, Wang, Wenjie, Zhao, Xiaoyan, Feng, Fuli, He, Xiangnan, Chua, Tat-Seng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917315805708288
author Wang, Chengbing
Zhang, Yang
Wang, Wenjie
Zhao, Xiaoyan
Feng, Fuli
He, Xiangnan
Chua, Tat-Seng
author_facet Wang, Chengbing
Zhang, Yang
Wang, Wenjie
Zhao, Xiaoyan
Feng, Fuli
He, Xiangnan
Chua, Tat-Seng
contents Preference alignment has enabled large language models (LLMs) to better reflect human expectations, but current methods mostly optimize for population-level preferences, overlooking individual users. Personalization is essential, yet early approaches-such as prompt customization or fine-tuning-struggle to reason over implicit preferences, limiting real-world effectiveness. Recent "think-then-generate" methods address this by reasoning before response generation. However, they face challenges in long-form generation: their static one-shot reasoning must capture all relevant information for the full response generation, making learning difficult and limiting adaptability to evolving content. To address this issue, we propose FlyThinker, an efficient "think-while-generating" framework for personalized long-form generation. FlyThinker employs a separate reasoning model that generates latent token-level reasoning in parallel, which is fused into the generation model to dynamically guide response generation. This design enables reasoning and generation to run concurrently, ensuring inference efficiency. In addition, the reasoning model is designed to depend only on previous responses rather than its own prior outputs, which preserves training parallelism across different positions-allowing all reasoning tokens for training data to be produced in a single forward pass like standard LLM training, ensuring training efficiency. Extensive experiments on real-world benchmarks demonstrate that FlyThinker achieves better personalized generation while keeping training and inference efficiency. Our code is available at https://github.com/wcb0219-sketch/FlyThinker.git.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06690
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
Wang, Chengbing
Zhang, Yang
Wang, Wenjie
Zhao, Xiaoyan
Feng, Fuli
He, Xiangnan
Chua, Tat-Seng
Computation and Language
Preference alignment has enabled large language models (LLMs) to better reflect human expectations, but current methods mostly optimize for population-level preferences, overlooking individual users. Personalization is essential, yet early approaches-such as prompt customization or fine-tuning-struggle to reason over implicit preferences, limiting real-world effectiveness. Recent "think-then-generate" methods address this by reasoning before response generation. However, they face challenges in long-form generation: their static one-shot reasoning must capture all relevant information for the full response generation, making learning difficult and limiting adaptability to evolving content. To address this issue, we propose FlyThinker, an efficient "think-while-generating" framework for personalized long-form generation. FlyThinker employs a separate reasoning model that generates latent token-level reasoning in parallel, which is fused into the generation model to dynamically guide response generation. This design enables reasoning and generation to run concurrently, ensuring inference efficiency. In addition, the reasoning model is designed to depend only on previous responses rather than its own prior outputs, which preserves training parallelism across different positions-allowing all reasoning tokens for training data to be produced in a single forward pass like standard LLM training, ensuring training efficiency. Extensive experiments on real-world benchmarks demonstrate that FlyThinker achieves better personalized generation while keeping training and inference efficiency. Our code is available at https://github.com/wcb0219-sketch/FlyThinker.git.
title Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
topic Computation and Language
url https://arxiv.org/abs/2512.06690