BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ou, Litu, Li, Kuan, Yin, Huifeng, Zhang, Liwen, Zhang, Zhongwang, Wu, Xixi, Ye, Rui, Qiao, Zile, Xie, Pengjun, Zhou, Jingren, Jiang, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911236015259648
author Ou, Litu
Li, Kuan
Yin, Huifeng
Zhang, Liwen
Zhang, Zhongwang
Wu, Xixi
Ye, Rui
Qiao, Zile
Xie, Pengjun
Zhou, Jingren
Jiang, Yong
author_facet Ou, Litu
Li, Kuan
Yin, Huifeng
Zhang, Liwen
Zhang, Zhongwang
Wu, Xixi
Ye, Rui
Qiao, Zile
Xie, Pengjun
Zhou, Jingren
Jiang, Yong
contents Confidence in LLMs is a useful indicator of model uncertainty and answer reliability. Existing work mainly focused on single-turn scenarios, while research on confidence in complex multi-turn interactions is limited. In this paper, we investigate whether LLM-based search agents have the ability to communicate their own confidence through verbalized confidence scores after long sequences of actions, a significantly more challenging task compared to outputting confidence in a single interaction. Experimenting on open-source agentic models, we first find that models exhibit much higher task accuracy at high confidence while having near-zero accuracy when confidence is low. Based on this observation, we propose Test-Time Scaling (TTS) methods that use confidence scores to determine answer quality, encourage the model to try again until reaching a satisfactory confidence level. Results show that our proposed methods significantly reduce token consumption while demonstrating competitive performance compared to baseline fixed budget TTS methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23458
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents
Ou, Litu
Li, Kuan
Yin, Huifeng
Zhang, Liwen
Zhang, Zhongwang
Wu, Xixi
Ye, Rui
Qiao, Zile
Xie, Pengjun
Zhou, Jingren
Jiang, Yong
Computation and Language
Artificial Intelligence
Confidence in LLMs is a useful indicator of model uncertainty and answer reliability. Existing work mainly focused on single-turn scenarios, while research on confidence in complex multi-turn interactions is limited. In this paper, we investigate whether LLM-based search agents have the ability to communicate their own confidence through verbalized confidence scores after long sequences of actions, a significantly more challenging task compared to outputting confidence in a single interaction. Experimenting on open-source agentic models, we first find that models exhibit much higher task accuracy at high confidence while having near-zero accuracy when confidence is low. Based on this observation, we propose Test-Time Scaling (TTS) methods that use confidence scores to determine answer quality, encourage the model to try again until reaching a satisfactory confidence level. Results show that our proposed methods significantly reduce token consumption while demonstrating competitive performance compared to baseline fixed budget TTS methods.
title BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.23458