In-Context Optimization for Retrieval-Augmented Generation: A Gradient-Descent Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Mingchen, Huang, Jiatan, Zhang, Chuxu, Zhao, Liang, Yu, Hong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918523250409472
author Li, Mingchen
Huang, Jiatan
Zhang, Chuxu
Zhao, Liang
Yu, Hong
author_facet Li, Mingchen
Huang, Jiatan
Zhang, Chuxu
Zhao, Liang
Yu, Hong
contents In-context learning has recently been linked to implicit gradient descent in linear self-attention models, suggesting that context can induce a forward-pass update. Retrieval-augmented generation (RAG) also relies on context, but retrieved documents are usually treated as static evidence rather than signals for adaptation. We study RAG as an in-context optimization process. First, we show that one linear self-attention layer can implement one gradient-descent step on a unified linearized RAG objective covering both projection-based and dot-product retrieval interfaces. This gives an exact regime where retrieval-augmented prediction and in-context optimization coincide. We use this result not as a literal model of LLM computation, but as a guide for adapting the interaction between queries and retrieved evidence. We then test the boundary of this correspondence: it remains stable under controlled linear extensions, but becomes feature-distribution dependent under nonlinear architectures. Finally, we turn this view into a lightweight method for frozen RAG LLMs. The method keeps the retriever and backbone fixed, and predicts a context-conditioned update to a generator-side evidence-use interface. Across seven QA benchmarks, two retrievers, and two frozen LLM backbones, this forward-only update improves a shared-interface baseline, transfers to held-out tasks, and approaches test-time gradient adaptation at much lower per-query cost.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26356
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle In-Context Optimization for Retrieval-Augmented Generation: A Gradient-Descent Perspective
Li, Mingchen
Huang, Jiatan
Zhang, Chuxu
Zhao, Liang
Yu, Hong
Computation and Language
In-context learning has recently been linked to implicit gradient descent in linear self-attention models, suggesting that context can induce a forward-pass update. Retrieval-augmented generation (RAG) also relies on context, but retrieved documents are usually treated as static evidence rather than signals for adaptation. We study RAG as an in-context optimization process. First, we show that one linear self-attention layer can implement one gradient-descent step on a unified linearized RAG objective covering both projection-based and dot-product retrieval interfaces. This gives an exact regime where retrieval-augmented prediction and in-context optimization coincide. We use this result not as a literal model of LLM computation, but as a guide for adapting the interaction between queries and retrieved evidence. We then test the boundary of this correspondence: it remains stable under controlled linear extensions, but becomes feature-distribution dependent under nonlinear architectures. Finally, we turn this view into a lightweight method for frozen RAG LLMs. The method keeps the retriever and backbone fixed, and predicts a context-conditioned update to a generator-side evidence-use interface. Across seven QA benchmarks, two retrievers, and two frozen LLM backbones, this forward-only update improves a shared-interface baseline, transfers to held-out tasks, and approaches test-time gradient adaptation at much lower per-query cost.
title In-Context Optimization for Retrieval-Augmented Generation: A Gradient-Descent Perspective
topic Computation and Language
url https://arxiv.org/abs/2605.26356