Saved in:
Bibliographic Details
Main Author: Franko, Uria
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2602.17046
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917281542438912
author Franko, Uria
author_facet Franko, Uria
contents Large Language Model (LLM) agents often run for many steps while re-ingesting long system instructions and large tool catalogs each turn. This increases cost, agent derailment probability, latency, and tool-selection errors. We propose Instruction-Tool Retrieval (ITR), a RAG variant that retrieves, per step, only the minimal system-prompt fragments and the smallest necessary subset of tools. ITR composes a dynamic runtime system prompt and exposes a narrowed toolset with confidence-gated fallbacks. Using a controlled benchmark with internally consistent numbers, ITR reduces per-step context tokens by 95%, improves correct tool routing by 32% relative, and cuts end-to-end episode cost by 70% versus a monolithic baseline. These savings enable agents to run 2-20x more loops within context limits. Savings compound with the number of agent steps, making ITR particularly valuable for long-running autonomous agents. We detail the method, evaluation protocol, ablations, and operational guidance for practical deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2602_17046
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic System Instructions and Tool Exposure for Efficient Agentic LLMs
Franko, Uria
Artificial Intelligence
Large Language Model (LLM) agents often run for many steps while re-ingesting long system instructions and large tool catalogs each turn. This increases cost, agent derailment probability, latency, and tool-selection errors. We propose Instruction-Tool Retrieval (ITR), a RAG variant that retrieves, per step, only the minimal system-prompt fragments and the smallest necessary subset of tools. ITR composes a dynamic runtime system prompt and exposes a narrowed toolset with confidence-gated fallbacks. Using a controlled benchmark with internally consistent numbers, ITR reduces per-step context tokens by 95%, improves correct tool routing by 32% relative, and cuts end-to-end episode cost by 70% versus a monolithic baseline. These savings enable agents to run 2-20x more loops within context limits. Savings compound with the number of agent steps, making ITR particularly valuable for long-running autonomous agents. We detail the method, evaluation protocol, ablations, and operational guidance for practical deployment.
title Dynamic System Instructions and Tool Exposure for Efficient Agentic LLMs
topic Artificial Intelligence
url https://arxiv.org/abs/2602.17046