SIA: Self Improving AI with Harness & Weight Updates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hebbar, Prannay, Manawat, Yogendra, Verboomen, Samuel, Ivanova, Alesia, Palanimalai, Selvam, Bhatia, Kunal, Baskaran, Vignesh
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917542045417472
author Hebbar, Prannay
Manawat, Yogendra
Verboomen, Samuel
Ivanova, Alesia
Palanimalai, Selvam
Bhatia, Kunal
Baskaran, Vignesh
author_facet Hebbar, Prannay
Manawat, Yogendra
Verboomen, Samuel
Ivanova, Alesia
Palanimalai, Selvam
Bhatia, Kunal
Baskaran, Vignesh
contents Humans are the bottleneck in building and improving AI. Both the models and the agents that wrap them are written, tuned, and corrected by people. The long-horizon goal of an AI that can figure out how to improve itself remains open. Two largely disjoint research lines attack this bottleneck. The harness-update school has a meta-agent rewrite the scaffold of a task-specific agent (its tools, prompts, retry logic, and search procedure) while the model weights are held fixed. The test-time training school uses hand-written RL pipelines to update the model's own weights on task feedback while the harness is held fixed. These two silos operate in isolation. We propose SIA, a self-improving loop in which a language-model agent (the Feedback-Agent) updates both the harness and the weights of a task-specific agent. We evaluate across three contrasting domains: Chinese legal charge classification, low-level GPU kernel optimisation, and single-cell RNA denoising. Combining both levers outperforms scaffold iteration alone on all three benchmarks. SIA-W+H achieves 25.1% over prior SOTA on LawBench, 12.4% faster GPU kernels than prior SOTA (1,017 vs 1,161 μs), and 20.4% over prior SOTA on denoising. Harness updates make the model agentic, shaping how it searches and acts, while weight updates build the domain intuition that no prompt or scaffold can instil.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27276
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SIA: Self Improving AI with Harness & Weight Updates
Hebbar, Prannay
Manawat, Yogendra
Verboomen, Samuel
Ivanova, Alesia
Palanimalai, Selvam
Bhatia, Kunal
Baskaran, Vignesh
Artificial Intelligence
Computation and Language
I.2.6; I.2.11; I.2.8
Humans are the bottleneck in building and improving AI. Both the models and the agents that wrap them are written, tuned, and corrected by people. The long-horizon goal of an AI that can figure out how to improve itself remains open. Two largely disjoint research lines attack this bottleneck. The harness-update school has a meta-agent rewrite the scaffold of a task-specific agent (its tools, prompts, retry logic, and search procedure) while the model weights are held fixed. The test-time training school uses hand-written RL pipelines to update the model's own weights on task feedback while the harness is held fixed. These two silos operate in isolation. We propose SIA, a self-improving loop in which a language-model agent (the Feedback-Agent) updates both the harness and the weights of a task-specific agent. We evaluate across three contrasting domains: Chinese legal charge classification, low-level GPU kernel optimisation, and single-cell RNA denoising. Combining both levers outperforms scaffold iteration alone on all three benchmarks. SIA-W+H achieves 25.1% over prior SOTA on LawBench, 12.4% faster GPU kernels than prior SOTA (1,017 vs 1,161 μs), and 20.4% over prior SOTA on denoising. Harness updates make the model agentic, shaping how it searches and acts, while weight updates build the domain intuition that no prompt or scaffold can instil.
title SIA: Self Improving AI with Harness & Weight Updates
topic Artificial Intelligence
Computation and Language
I.2.6; I.2.11; I.2.8
url https://arxiv.org/abs/2605.27276