O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mei, Jianbiao, Hu, Tao, Fu, Daocheng, Wen, Licheng, Yang, Xuemeng, Wu, Rong, Cai, Pinlong, Cai, Xinyu, Gao, Xing, Yang, Yu, Xie, Chengjun, Shi, Botian, Liu, Yong, Qiao, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916758737125376
author Mei, Jianbiao
Hu, Tao
Fu, Daocheng
Wen, Licheng
Yang, Xuemeng
Wu, Rong
Cai, Pinlong
Cai, Xinyu
Gao, Xing
Yang, Yu
Xie, Chengjun
Shi, Botian
Liu, Yong
Qiao, Yu
author_facet Mei, Jianbiao
Hu, Tao
Fu, Daocheng
Wen, Licheng
Yang, Xuemeng
Wu, Rong
Cai, Pinlong
Cai, Xinyu
Gao, Xing
Yang, Yu
Xie, Chengjun
Shi, Botian
Liu, Yong
Qiao, Yu
contents Large Language Models (LLMs), despite their advancements, are fundamentally limited by their static parametric knowledge, hindering performance on tasks requiring open-domain up-to-date information. While enabling LLMs to interact with external knowledge environments is a promising solution, current efforts primarily address closed-end problems. Open-ended questions, which characterized by lacking a standard answer or providing non-unique and diverse answers, remain underexplored. To bridge this gap, we present O$^2$-Searcher, a novel search agent leveraging reinforcement learning to effectively tackle both open-ended and closed-ended questions in the open domain. O$^2$-Searcher leverages an efficient, locally simulated search environment for dynamic knowledge acquisition, effectively decoupling the external world knowledge from model's sophisticated reasoning processes. It employs a unified training mechanism with meticulously designed reward functions, enabling the agent to identify problem types and adapt different answer generation strategies. Furthermore, to evaluate performance on complex open-ended tasks, we construct O$^2$-QA, a high-quality benchmark featuring 300 manually curated, multi-domain open-ended questions with associated web page caches. Extensive experiments show that O$^2$-Searcher, using only a 3B model, significantly surpasses leading LLM agents on O$^2$-QA. It also achieves SOTA results on various closed-ended QA benchmarks against similarly-sized models, while performing on par with much larger ones.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16582
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering
Mei, Jianbiao
Hu, Tao
Fu, Daocheng
Wen, Licheng
Yang, Xuemeng
Wu, Rong
Cai, Pinlong
Cai, Xinyu
Gao, Xing
Yang, Yu
Xie, Chengjun
Shi, Botian
Liu, Yong
Qiao, Yu
Computation and Language
Artificial Intelligence
Large Language Models (LLMs), despite their advancements, are fundamentally limited by their static parametric knowledge, hindering performance on tasks requiring open-domain up-to-date information. While enabling LLMs to interact with external knowledge environments is a promising solution, current efforts primarily address closed-end problems. Open-ended questions, which characterized by lacking a standard answer or providing non-unique and diverse answers, remain underexplored. To bridge this gap, we present O$^2$-Searcher, a novel search agent leveraging reinforcement learning to effectively tackle both open-ended and closed-ended questions in the open domain. O$^2$-Searcher leverages an efficient, locally simulated search environment for dynamic knowledge acquisition, effectively decoupling the external world knowledge from model's sophisticated reasoning processes. It employs a unified training mechanism with meticulously designed reward functions, enabling the agent to identify problem types and adapt different answer generation strategies. Furthermore, to evaluate performance on complex open-ended tasks, we construct O$^2$-QA, a high-quality benchmark featuring 300 manually curated, multi-domain open-ended questions with associated web page caches. Extensive experiments show that O$^2$-Searcher, using only a 3B model, significantly surpasses leading LLM agents on O$^2$-QA. It also achieves SOTA results on various closed-ended QA benchmarks against similarly-sized models, while performing on par with much larger ones.
title O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.16582