Google Cloud AI Research, with UNC-Chapel Hill, Stanford and Washington University in St. Louis, has released RRSI (Regularized Recursive Self-Improvement). It lets an LLM agent rewrite its own harness: prompts, tools, memory, control flow and sub-agents. Model weights never change. RRSI constrains the improvement loop itself, so gains hold on benchmarks the agent never optimized against.
Deployable? Yes, as a research framework. The code is Apache 2.0, needs Python 3.10+, and accepts any LiteLLM model string. Defaults assume Claude Opus 4.8 on Vertex AI.
Why Self-Improving Harnesses Overfit
Harness evolution loops propose edits, score them on a fixed evolve set and keep the winner. The same tasks are reused every round, so the loop can memorize them. The RRSI research names 3 failure modes: benchmark-specific fitting, noise chasing and complexity accumulation. Each one widens the gap between evolve-set scores and real transfer.
How RRSI Works
RRSI keeps every harness component editable. It regularizes how the search moves instead.
Proposal side
Annealed edit budget: a cosine schedule lets early rounds bundle several edits. Late rounds allow a single attributable change.
Evidence-aware credit: each candidate is logged with its component, hypothesis, diff, score change and cost change. The proposer reads this ledger, so falsified ideas are not retried.
Structured exploration: when progress stalls inside the noise band, budget shifts to components the run never touched.
Selection side
Leakage critic: rejects task names, entities, answers or benchmark-specific logic before any scoring.
Noise-adjusted floor: gains must clear the variance measured on the unchanged base harness.
Cost rule: extra inference tokens must be paid for by measured gain.
Pruning: components that stop producing gains become deletion targets.
The research team frame these as analogies to classic regularizers. The edit budget maps to L0, pruning to Lasso (L1) and the cost rule to Ridge (L2).
Results Across 8 Benchmarks
Terminal-Bench 2.1 (evolve split): 74.2% to 80.2%.
SWE-bench Verified (never used for selection): 82.0% to 83.8%.
Out of distribution: JobBench +4.7, GDPval +3.5 and APEX-Agents +3.7 points.
EngDesign (evolve) +4.9; Frontier-Eng +4.3 Medal points.
Harvey LAB: +1.1 on the evolve split, +2.3 on its held-out split.
All 6 held-out splits improved. With Gemini 3.5 Flash as the policy, Terminal-Bench 2.1 rose from 64.6 to 78.7. SWE-bench Verified rose from 76.8 to 79.0.
The harness is also lighter. On the agentic workspace instance, RRSI uses 2.42M policy tokens per trial. Unregularized evolution uses 3.80M. The abstract reports this as 30% fewer; the project page says 36%.
RRSI vs Closest Competitors
Scores come from Table 1 of the RRSI research paper. All methods share the same starting harness, policy, evolve split and candidate budget.
Feature
RRSI
Meta-Harness
AHE
TTHE
HarnessX
Core idea
Regularized proposal and selection
Agentic proposer over code, scores and traces of all prior candidates
Observability-driven loop; edits paired with verified predictions
Evolves harness during test time, no gold labels
Modular typed primitives, trace-driven adaptation
Model weights
Frozen
Frozen
Frozen
Frozen
Frozen
Cost rule and pruning
Yes
No*
No*
No*
No*
Harvey LAB evolve score
90.5
93.0
90.7
91.1
91.8
OOD average (H0 = 39.7)
43.6
40.6
39.2
38.0
39.7
*Per the RRSI research team. OOD average is the mean of JobBench, GDPval and APEX-Agents, computed from Table 1.
Meta-Harness leads on the Harvey LAB evolve split. RRSI has the smallest evolve gain but the only OOD average more than 1 point above H0.
Interactive Explainer
How RRSI Regularizes Agent Self-Improvement
The model stays frozen. The harness (prompts, tools, memory, control flow) evolves, but every edit must pass the regularizers.
1. Run a round
2. Edit budget
3. Acceptance gate
4. Results
Pick a candidate edit, then press Run. Watch where RRSI stops it.
Generic tool fix
Hardcoded task name
Gain inside noise
Token-hungry gain
Stale component
▶ Run
ProposerReads full edit ledger, shrinking edit budget
Leakage criticScreens diff before scoring
EvaluateRun on evolve set
GateNoise floor + cost rule
Harness Ht+1Accepted, then pruning check
Candidates are illustrative. The rules they hit are the ones described in the RRSI paper.
// edit ledger: component | hypothesis | Δscore | Δcost | verdict
bt = ⌈ bmin + (bmax − bmin) · ½(1 + cos(πt / T)) ⌉
bmax (edits per candidate, early) 5
bmin (edits per candidate, late) 1
T (rounds) 16
Early rounds may bundle coordinated edits to find a mechanism. Late rounds get single, attributable changes.
round 0round T−1
Cosine schedule from the paper (Eq. 4). Slider values are for exploration; the paper lists its own settings in Appendix D.
ΔS, score gain vs incumbent (pts) 2.0
ΔC, token cost change (%) 20
δ, noise band from repeated base runs (pts) 1.5
β0 + β1ΔS cost allowance: β1 (% per pt) 8
?Noise-adjusted floor: score must not drop more than δ below the best so far
?Clear gain: ΔS > δ, otherwise the paper’s within-band rule applies
?Cost rule: ΔC ≤ β0 + β1ΔS
…
Illustrative thresholds (β0 fixed at 5% here). In RRSI, δ is estimated on the unchanged base harness and β values are tuned on the evolve set, then frozen.
All 8 benchmarks
vs prior methods
Token cost
Data: RRSI paper · GitHub · project pageBuilt by Marktechpost
‘});
var B=$(‘#rr-bars’);B.innerHTML=h;setTimeout(function(){B.querySelectorAll(‘div’).forEach(function(d){d.style.height=d.getAttribute(‘data-h’)+’%’})},30)}
[‘#s-bmax’,’#s-bmin’,’#s-T’].forEach(function(s){$(s).oninput=budget});budget();
/* PANE 3 */
function setChk(id,st){var e=$(id),ic=e.querySelector(‘.ic’);var col=st===1?’#34A853′:st===0?’#EA4335′:’#5f6368′;ic.style.background=col;ic.textContent=st===1?’u2713′:st===0?’u2715′:’-‘;e.style.borderColor=col}
function gate(){var ds=+$(‘#s-ds’).value,dc=+$(‘#s-dc’).value,d=+$(‘#s-d’).value,b1=+$(‘#s-b1’).value,b0=5;
$(‘#v-ds’).textContent=ds.toFixed(1);$(‘#v-dc’).textContent=dc;$(‘#v-d’).textContent=d.toFixed(1);$(‘#v-b1’).textContent=b1;
var allow=b0+b1*ds;$(‘#c3t’).innerHTML=’Cost rule: ΔC ≤ β0 + β1ΔS = ‘+allow.toFixed(1)+’%’;
var f=ds>=-d,clear=ds>d,cost=dc<=allow,V=$(‘#rr-verdict’),txt,col;
setChk(‘#c1’,f?1:0);
if(!f){setChk(‘#c2’,-1);setChk(‘#c3′,-1);txt=’REJECT: fell below the noise-adjusted floor’;col=’#EA4335′}
else if(!clear){setChk(‘#c2’,0);setChk(‘#c3′,-1);txt=’WITHIN NOISE: not credited as a clear gain’;col=’#FBBC05′}
else{setChk(‘#c2’,1);setChk(‘#c3′,cost?1:0);txt=cost?’ACCEPT: real gain that pays for its tokens’:’REJECT: gain does not pay for extra tokens’;col=cost?’#34A853′:’#EA4335′}
V.textContent=txt;V.style.background=col+’22’;V.style.color=col;V.style.border=’1px solid ‘+col}
[‘#s-ds’,’#s-dc’,’#s-d’,’#s-b1′].forEach(function(s){$(s).oninput=gate});gate();
/* PANE 4 */
var D=[
{t:’H0 vs RRSI, frozen Claude Opus 4.8 (evolve split marked)’,max:100,u:”,rows:[
[‘Terminal-Bench 2.1′,’evolve’,74.2,80.2],[‘SWE-bench Verified’,’OOD’,82.0,83.8],[‘Harvey LAB’,’evolve’,89.4,90.5],[‘Harvey LAB held-out’,’ID held-out’,86.9,89.2],[‘JobBench’,’OOD’,36.0,40.7],[‘GDPval’,’OOD’,48.8,52.3],[‘APEX-Agents’,’OOD’,34.2,37.9],[‘EngDesign’,’evolve’,50.0,54.9],[‘Frontier-Eng’,’OOD’,17.7,22.0]],
n:’Grey bar = unevolved harness H0, blue bar = RRSI. Frontier-Eng is Medal points; GDPval is win rate vs human experts; Harvey LAB is rubric criteria passed.’},
{t:’Out-of-distribution average (JobBench, GDPval, APEX-Agents)’,max:50,u:”,single:1,rows:[[‘H0 (no evolution)’,”,39.7],[‘Meta-Harness’,”,40.6],[‘AHE’,”,39.2],[‘TTHE’,”,38.0],[‘HarnessX’,”,39.7],[‘RRSI’,”,43.6]],
n:’Averages computed from Table 1 of the RRSI paper. All methods start from the same H0, policy, evolve split and candidate budget.’},
{t:’Policy tokens per trial (millions), agentic workspace’,max:4,u:’M’,single:1,rows:[[‘H0 (no evolution)’,’OOD 39.7′,1.56],[‘Unregularized evolution’,’OOD 40.3′,3.80],[‘w/o acceptance regularizers’,’OOD 41.0′,3.59],[‘w/o proposal regularizers’,’OOD 41.9′,2.69],[‘RRSI’,’OOD 43.6′,2.42]],
n:’From the ablation table: RRSI uses 2.42M tokens per trial vs 3.80M for unregularized evolution, about 36% fewer, with the best OOD average.’}],view=0;
function chart(v){var d=D[v],h=’
‘+d.t+’
‘;
d.rows.forEach(function(r){var isR=r[0]===’RRSI’;
if(d.single){h+=’
‘+r[0]+(r[1]?”+r[1]+”:”)+”+r[2].toFixed(d.u?2:1)+d.u+’
‘}
else{h+=’
‘+r[0]+”+r[1]+”+r[2].toFixed(1)+’ → ‘+r[3].toFixed(1)+’ (+’+(r[3]-r[2]).toFixed(1)+’)
‘}});
$(‘#rr-chart’).innerHTML=h;$(‘#rr-cnote’).textContent=d.n;
setTimeout(function(){$$(‘#rr-chart i’).forEach(function(i){i.style.width=i.getAttribute(‘data-w’)+’%’});resize()},40)}
$$(‘#rr-views .btn’).forEach(function(b){b.onclick=function(){$$(‘#rr-views .btn’).forEach(function(x){x.classList.remove(‘on’)});b.classList.add(‘on’);view=+b.getAttribute(‘data-v’);chart(view)}});
chart(0);
window.addEventListener(‘load’,resize);window.addEventListener(‘resize’,resize);setTimeout(resize,300);
})();