Search Arena: Analyzing Search-Augmented LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miroyan, Mihran, Wu, Tsung-Han, King, Logan, Li, Tianle, Pan, Jiayi, Hu, Xinyan, Chiang, Wei-Lin, Angelopoulos, Anastasios N., Darrell, Trevor, Norouzi, Narges, Gonzalez, Joseph E.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908860765175808
author Miroyan, Mihran
Wu, Tsung-Han
King, Logan
Li, Tianle
Pan, Jiayi
Hu, Xinyan
Chiang, Wei-Lin
Angelopoulos, Anastasios N.
Darrell, Trevor
Norouzi, Narges
Gonzalez, Joseph E.
author_facet Miroyan, Mihran
Wu, Tsung-Han
King, Logan
Li, Tianle
Pan, Jiayi
Hu, Xinyan
Chiang, Wei-Lin
Angelopoulos, Anastasios N.
Darrell, Trevor
Norouzi, Narges
Gonzalez, Joseph E.
contents Search-augmented language models combine web search with Large Language Models (LLMs) to improve response groundedness and freshness. However, analyzing these systems remains challenging: existing datasets are limited in scale and narrow in scope, often constrained to static, single-turn, fact-checking questions. In this work, we introduce Search Arena, a crowd-sourced, large-scale, human-preference dataset of over 24,000 paired multi-turn user interactions with search-augmented LLMs. The dataset spans diverse intents and languages, and contains full system traces with around 12,000 human preference votes. Our analysis reveals that user preferences are influenced by the number of citations, even when the cited content does not directly support the attributed claims, uncovering a gap between perceived and actual credibility. Furthermore, user preferences vary across cited sources, revealing that community-driven platforms are generally preferred and static encyclopedic sources are not always appropriate and reliable. To assess performance across different settings, we conduct cross-arena analyses by testing search-augmented LLMs in a general-purpose chat environment and conventional LLMs in search-intensive settings. We find that web search does not degrade and may even improve performance in non-search settings; however, the quality in search settings is significantly affected if solely relying on the model's parametric knowledge. We open-sourced the dataset to support future research. Our dataset and code are available at: https://github.com/lmarena/search-arena.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05334
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Search Arena: Analyzing Search-Augmented LLMs
Miroyan, Mihran
Wu, Tsung-Han
King, Logan
Li, Tianle
Pan, Jiayi
Hu, Xinyan
Chiang, Wei-Lin
Angelopoulos, Anastasios N.
Darrell, Trevor
Norouzi, Narges
Gonzalez, Joseph E.
Computation and Language
Information Retrieval
Machine Learning
Search-augmented language models combine web search with Large Language Models (LLMs) to improve response groundedness and freshness. However, analyzing these systems remains challenging: existing datasets are limited in scale and narrow in scope, often constrained to static, single-turn, fact-checking questions. In this work, we introduce Search Arena, a crowd-sourced, large-scale, human-preference dataset of over 24,000 paired multi-turn user interactions with search-augmented LLMs. The dataset spans diverse intents and languages, and contains full system traces with around 12,000 human preference votes. Our analysis reveals that user preferences are influenced by the number of citations, even when the cited content does not directly support the attributed claims, uncovering a gap between perceived and actual credibility. Furthermore, user preferences vary across cited sources, revealing that community-driven platforms are generally preferred and static encyclopedic sources are not always appropriate and reliable. To assess performance across different settings, we conduct cross-arena analyses by testing search-augmented LLMs in a general-purpose chat environment and conventional LLMs in search-intensive settings. We find that web search does not degrade and may even improve performance in non-search settings; however, the quality in search settings is significantly affected if solely relying on the model's parametric knowledge. We open-sourced the dataset to support future research. Our dataset and code are available at: https://github.com/lmarena/search-arena.
title Search Arena: Analyzing Search-Augmented LLMs
topic Computation and Language
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2506.05334