Google RRSI Keeps Self-Improving AI Agents From Overfitting
AI agents that rewrite their own scaffolding tend to get very good at their test tasks and not much better at anything else. A new paper from Google Cloud AI Research and several universities describes a method that limits this effect and also cuts compute costs.
The harness does much of the work
Today's agents usually pair a fixed language model with a "harness." That is the layer of prompts, workflows, tools, memory and control logic that decides what the model sees at each step. It determines whether the agent opens the right file before editing it, whether it recovers after a mistake, and whether it hands back clean results.
The paper argues that much of the recent progress in agents has come from improving harnesses, not from new models. Projects like BootLoops show how far a well-built harness can take an existing model.
Harness tuning used to be manual. Engineers read through failed runs and patched the logic by hand. Newer approaches automate this. A language model rewrites the harness repeatedly, based on feedback from test tasks. The researchers describe this as a practical form of recursive self-improvement: the system generates feedback, uses it to change the harness, and the harness then shapes how the system behaves.
The memorization problem
The catch is overfitting. When an agent keeps optimizing against the same small set of tasks, it starts to memorize them. Training scores rise. Gains on unseen tasks shrink or vanish.
The paper names three ways this happens:
- The search picks up patterns that only work on one specific benchmark.
- It favors candidates that scored well by chance.
- It adds complexity that lifts the test score without making the agent more capable.
How RRSI works
The method is called RRSI, short for Regularized Recursive Self-Improvement of Agent Harnesses. The harness stays fully editable. RRSI adds controls at two points in the loop.
Proposing changes. A budget limits how many independent edits a candidate can bundle together. The budget shrinks over time. Early rounds allow larger rewrites. Later rounds only permit small changes whose effect can be clearly traced. The system also remembers past attempts so it does not keep retrying failed ideas. If progress stalls, it deliberately tries changes in parts of the harness it has not touched yet.
Accepting changes. A critic checks each proposal and rejects any that hardcode task names, solutions or other benchmark-specific shortcuts. Higher compute costs are only accepted if they come with a measurable performance gain. Components that stop helping are removed.
The results
The team tested RRSI on eight benchmarks covering coding, agentic office work and engineering design. The underlying model, Claude Opus 4.8, was frozen throughout. RRSI was compared with the unmodified baseline harness and four recent optimization methods.
According to the paper, RRSI gained up to 14.1 points on its training tasks and up to 4.7 points on five benchmarks it never saw. The largest unseen gain was on JobBench. It used about 30 percent fewer tokens at runtime than the unregularized version. On none of the unseen benchmarks did it fall below the baseline.
The comparison is the interesting part. Every method improved on training tasks. On new tasks, the picture flipped. Two methods ended up below the baseline harness. RRSI had the smallest training gain of all variants but was the only one clearly above baseline on unseen tasks. That trade-off is what the guardrails are designed to produce. Among the optimized harnesses, RRSI used the fewest tokens and steps, although the untouched baseline was leaner still.
Harnesses also transferred across models. A coding harness optimized with Gemini 3.5 Flash lifted the much weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points with no changes. The authors say the mechanisms found do not depend on the strength of the model used to discover them.
The study only covers harnesses around frozen models. Cases where model weights change are not addressed. The code is on GitHub.
Why It Matters
For teams building agents, this suggests the harness deserves the same scrutiny as the model, including the risk of tuning it into a benchmark specialist. The paper points to a familiar warning: on ARC-AGI-3, Opus 4.6 with a purpose-built harness scored 97.1 percent in a familiar environment and 0 percent in an unfamiliar one.
RRSI also fits a wider trend of improving agents without retraining models. Nvidia's SoL-Pi has a research agent rebuild coding-agent harnesses, cutting token use by up to 49 percent. Google earlier had agents "dream" about past search runs. Developers already customize agent behavior by hand. Automating that safely is the next step.
The critic is the piece to watch. Agents are not always good at judging their own output. It is worth watching whether independent groups reproduce these unseen-task gains, and whether the approach holds up once model weights also change.
