Natural Language Query to Configuration for Retrieval Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Melissa Z., Arabzadeh, Negar, Jacob, Mathew, Kazhamiaka, Fiodar, Choukse, Esha, Zaharia, Matei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913164653756416
author Pan, Melissa Z.
Arabzadeh, Negar
Jacob, Mathew
Kazhamiaka, Fiodar
Choukse, Esha
Zaharia, Matei
author_facet Pan, Melissa Z.
Arabzadeh, Negar
Jacob, Mathew
Kazhamiaka, Fiodar
Choukse, Esha
Zaharia, Matei
contents Modern retrieval agents expose many configuration choices -- LLM, retriever, number of documents, number of hops, and synthesis strategy -- each shaping both answer quality and serving cost. Today, these pipelines are typically hand-tuned once per workload, leaving substantial per-query optimization untapped. We formulate the problem: given a natural-language query and either an accuracy or a budget target, select from a predefined pipeline catalog the configuration that minimizes cost or maximizes accuracy at inference time. We propose **BRANE**, which uses an LLM to convert each query into workload-specific characteristics, then trains a lightweight per-configuration predictor that estimates whether the pipeline will answer the query correctly. At inference time, **BRANE** selects the configuration that maximizes predicted correctness penalized by cost, exposing a tunable cost-quality tradeoff without retraining. Across MuSiQue, BrowseComp-Plus, and FinanceBench, **BRANE** consistently pushes the cost-quality Pareto frontier, matches the best fixed configuration's accuracy at up to 89% lower cost, and outperforms LLM-routing, rule-based, and fine-tuned Qwen3-4B baselines. These results show that per-query configuration of the full retrieval pipeline is a practical alternative to static workload-level tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27361
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Natural Language Query to Configuration for Retrieval Agents
Pan, Melissa Z.
Arabzadeh, Negar
Jacob, Mathew
Kazhamiaka, Fiodar
Choukse, Esha
Zaharia, Matei
Artificial Intelligence
Systems and Control
Modern retrieval agents expose many configuration choices -- LLM, retriever, number of documents, number of hops, and synthesis strategy -- each shaping both answer quality and serving cost. Today, these pipelines are typically hand-tuned once per workload, leaving substantial per-query optimization untapped. We formulate the problem: given a natural-language query and either an accuracy or a budget target, select from a predefined pipeline catalog the configuration that minimizes cost or maximizes accuracy at inference time. We propose **BRANE**, which uses an LLM to convert each query into workload-specific characteristics, then trains a lightweight per-configuration predictor that estimates whether the pipeline will answer the query correctly. At inference time, **BRANE** selects the configuration that maximizes predicted correctness penalized by cost, exposing a tunable cost-quality tradeoff without retraining. Across MuSiQue, BrowseComp-Plus, and FinanceBench, **BRANE** consistently pushes the cost-quality Pareto frontier, matches the best fixed configuration's accuracy at up to 89% lower cost, and outperforms LLM-routing, rule-based, and fine-tuned Qwen3-4B baselines. These results show that per-query configuration of the full retrieval pipeline is a practical alternative to static workload-level tuning.
title Natural Language Query to Configuration for Retrieval Agents
topic Artificial Intelligence
Systems and Control
url https://arxiv.org/abs/2605.27361