Leveraging Generative Models for Real-Time Query-Driven Text Summarization in Large-Scale Web Search

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xiong, Zeyu, Nan, Yixuan, Gao, Li, Tang, Hengzhu, Wang, Shuaiqiang, Wang, Junfeng, Yin, Dawei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916923179008000
author Xiong, Zeyu
Nan, Yixuan
Gao, Li
Tang, Hengzhu
Wang, Shuaiqiang
Wang, Junfeng
Yin, Dawei
author_facet Xiong, Zeyu
Nan, Yixuan
Gao, Li
Tang, Hengzhu
Wang, Shuaiqiang
Wang, Junfeng
Yin, Dawei
contents In the dynamic landscape of large-scale web search, Query-Driven Text Summarization (QDTS) aims to generate concise and informative summaries from textual documents based on a given query, which is essential for improving user engagement and facilitating rapid decision-making. Traditional extractive summarization models, based primarily on ranking candidate summary segments, have been the dominant approach in industrial applications. However, these approaches suffer from two key limitations: 1) The multi-stage pipeline often introduces cumulative information loss and architectural bottlenecks due to its weakest component; 2) Traditional models lack sufficient semantic understanding of both user queries and documents, particularly when dealing with complex search intents. In this study, we propose a novel framework to pioneer the application of generative models to address real-time QDTS in industrial web search. Our approach integrates large model distillation, supervised fine-tuning, direct preference optimization, and lookahead decoding to transform a lightweight model with only 0.1B parameters into a domain-specialized QDTS expert. Evaluated on multiple industry-relevant metrics, our model outperforms the production baseline and achieves a new state of the art. Furthermore, it demonstrates excellent deployment efficiency, requiring only 334 NVIDIA L20 GPUs to handle \textasciitilde50,000 queries per second under 55~ms average latency per query.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20559
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Generative Models for Real-Time Query-Driven Text Summarization in Large-Scale Web Search
Xiong, Zeyu
Nan, Yixuan
Gao, Li
Tang, Hengzhu
Wang, Shuaiqiang
Wang, Junfeng
Yin, Dawei
Computation and Language
Information Retrieval
In the dynamic landscape of large-scale web search, Query-Driven Text Summarization (QDTS) aims to generate concise and informative summaries from textual documents based on a given query, which is essential for improving user engagement and facilitating rapid decision-making. Traditional extractive summarization models, based primarily on ranking candidate summary segments, have been the dominant approach in industrial applications. However, these approaches suffer from two key limitations: 1) The multi-stage pipeline often introduces cumulative information loss and architectural bottlenecks due to its weakest component; 2) Traditional models lack sufficient semantic understanding of both user queries and documents, particularly when dealing with complex search intents. In this study, we propose a novel framework to pioneer the application of generative models to address real-time QDTS in industrial web search. Our approach integrates large model distillation, supervised fine-tuning, direct preference optimization, and lookahead decoding to transform a lightweight model with only 0.1B parameters into a domain-specialized QDTS expert. Evaluated on multiple industry-relevant metrics, our model outperforms the production baseline and achieves a new state of the art. Furthermore, it demonstrates excellent deployment efficiency, requiring only 334 NVIDIA L20 GPUs to handle \textasciitilde50,000 queries per second under 55~ms average latency per query.
title Leveraging Generative Models for Real-Time Query-Driven Text Summarization in Large-Scale Web Search
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2508.20559