GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Moll, Johannes, Corbeil, Jean-Philippe, Pan, Jiazhen, Hadamitzky, Martin, Rueckert, Daniel, Adams, Lisa, Bressem, Keno
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913170029805568
author Moll, Johannes
Corbeil, Jean-Philippe
Pan, Jiazhen
Hadamitzky, Martin
Rueckert, Daniel
Adams, Lisa
Bressem, Keno
author_facet Moll, Johannes
Corbeil, Jean-Philippe
Pan, Jiazhen
Hadamitzky, Martin
Rueckert, Daniel
Adams, Lisa
Bressem, Keno
contents LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment. Prior self-improvement methods accumulate natural-language guidance without checking that each new item preserves previously correct behavior, so a note that fixes one trajectory can silently regress another. We introduce GRASP (Gated Regression-Aware Skill Proposer), which treats agent improvement as a sequence of edits to a bounded skill library, admitting each candidate only if it produces a net improvement on a balanced held-out probe under a hard regression budget. We evaluate GRASP across five base models (gpt-oss-120b, DeepSeek V4 Flash, Gemini 3.1 Flash Lite, GPT-4.1, GPT-5.4) on two FHIR-based clinical benchmarks. On MedAgentBench, GRASP lifts gpt-oss-120b from 40.6% to 88.8%, exceeds the strongest of five self-improvement baselines by 21.0 points, and improves every other base model by 17.2 to 40.3 points. Ablations attribute the gain to comparative proposal generation, the acceptance gate, and the hard regression budget rather than to skill writing itself, which without validation is no better than using no skills. The mechanism generalizes beyond the clinical domain, improving agents on three of four non-clinical environments and remaining flat only where the action space is open-ended. Frozen libraries transfer across models, where skills from a stronger model improve weaker executors beyond what they learn for themselves while the reverse does not, an asymmetry that no ungated baseline reproduces.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29668
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents
Moll, Johannes
Corbeil, Jean-Philippe
Pan, Jiazhen
Hadamitzky, Martin
Rueckert, Daniel
Adams, Lisa
Bressem, Keno
Artificial Intelligence
Computation and Language
LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment. Prior self-improvement methods accumulate natural-language guidance without checking that each new item preserves previously correct behavior, so a note that fixes one trajectory can silently regress another. We introduce GRASP (Gated Regression-Aware Skill Proposer), which treats agent improvement as a sequence of edits to a bounded skill library, admitting each candidate only if it produces a net improvement on a balanced held-out probe under a hard regression budget. We evaluate GRASP across five base models (gpt-oss-120b, DeepSeek V4 Flash, Gemini 3.1 Flash Lite, GPT-4.1, GPT-5.4) on two FHIR-based clinical benchmarks. On MedAgentBench, GRASP lifts gpt-oss-120b from 40.6% to 88.8%, exceeds the strongest of five self-improvement baselines by 21.0 points, and improves every other base model by 17.2 to 40.3 points. Ablations attribute the gain to comparative proposal generation, the acceptance gate, and the hard regression budget rather than to skill writing itself, which without validation is no better than using no skills. The mechanism generalizes beyond the clinical domain, improving agents on three of four non-clinical environments and remaining flat only where the action space is open-ended. Frozen libraries transfer across models, where skills from a stronger model improve weaker executors beyond what they learn for themselves while the reverse does not, an asymmetry that no ungated baseline reproduces.
title GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.29668