Close Menu
InvesthenaInvesthena
    What's Hot

    Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting

    September 29, 2026

    Raydium’s Holders Revenue Reaches $4.43M as StonkFun Drives LaunchLab Volume

    September 29, 2026

    Stocks Settle Lower on Rising Bond Yields

    September 29, 2026
    InvesthenaInvesthena
    • Business
    • Economy
    • Investing
    • Stocks
    • Best Savings Accounts
    InvesthenaInvesthena
    Home»Business»Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
    Business

    Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting

    September 29, 2026
    Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
    Share
    Facebook Twitter LinkedIn Pinterest Email


    Google Cloud AI Research, with UNC-Chapel Hill, Stanford and Washington University in St. Louis, has released RRSI (Regularized Recursive Self-Improvement). It lets an LLM agent rewrite its own harness: prompts, tools, memory, control flow and sub-agents. Model weights never change. RRSI constrains the improvement loop itself, so gains hold on benchmarks the agent never optimized against.

    Deployable? Yes, as a research framework. The code is Apache 2.0, needs Python 3.10+, and accepts any LiteLLM model string. Defaults assume Claude Opus 4.8 on Vertex AI.

    Why Self-Improving Harnesses Overfit

    Harness evolution loops propose edits, score them on a fixed evolve set and keep the winner. The same tasks are reused every round, so the loop can memorize them. The RRSI research names 3 failure modes: benchmark-specific fitting, noise chasing and complexity accumulation. Each one widens the gap between evolve-set scores and real transfer.

    How RRSI Works

    RRSI keeps every harness component editable. It regularizes how the search moves instead.

    Proposal side

    • Annealed edit budget: a cosine schedule lets early rounds bundle several edits. Late rounds allow a single attributable change.
    • Evidence-aware credit: each candidate is logged with its component, hypothesis, diff, score change and cost change. The proposer reads this ledger, so falsified ideas are not retried.
    • Structured exploration: when progress stalls inside the noise band, budget shifts to components the run never touched.

    Selection side

    • Leakage critic: rejects task names, entities, answers or benchmark-specific logic before any scoring.
    • Noise-adjusted floor: gains must clear the variance measured on the unchanged base harness.
    • Cost rule: extra inference tokens must be paid for by measured gain.
    • Pruning: components that stop producing gains become deletion targets.

    The research team frame these as analogies to classic regularizers. The edit budget maps to L0, pruning to Lasso (L1) and the cost rule to Ridge (L2).

    Results Across 8 Benchmarks

    All 6 held-out splits improved. With Gemini 3.5 Flash as the policy, Terminal-Bench 2.1 rose from 64.6 to 78.7. SWE-bench Verified rose from 76.8 to 79.0.

    The harness is also lighter. On the agentic workspace instance, RRSI uses 2.42M policy tokens per trial. Unregularized evolution uses 3.80M. The abstract reports this as 30% fewer; the project page says 36%.

    RRSI vs Closest Competitors

    Scores come from Table 1 of the RRSI research paper. All methods share the same starting harness, policy, evolve split and candidate budget.

    *Per the RRSI research team. OOD average is the mean of JobBench, GDPval and APEX-Agents, computed from Table 1.

    Meta-Harness leads on the Harvey LAB evolve split. RRSI has the smallest evolve gain but the only OOD average more than 1 point above H0.

    Interactive Explainer

    How RRSI Regularizes Agent Self-Improvement

    The model stays frozen. The harness (prompts, tools, memory, control flow) evolves, but every edit must pass the regularizers.

    1. Run a round
    2. Edit budget
    3. Acceptance gate
    4. Results

    Pick a candidate edit, then press Run. Watch where RRSI stops it.

    Generic tool fix
    Hardcoded task name
    Gain inside noise
    Token-hungry gain
    Stale component
    ▶ Run

    ProposerReads full edit ledger, shrinking edit budget

    Leakage criticScreens diff before scoring

    EvaluateRun on evolve set

    GateNoise floor + cost rule

    Harness Ht+1Accepted, then pruning check

    Candidates are illustrative. The rules they hit are the ones described in the RRSI paper.

    // edit ledger: component | hypothesis | Δscore | Δcost | verdict

    All 8 benchmarks
    vs prior methods
    Token cost

    Getting Started

    git clone https://github.com/google-research/rrsi.git && cd rrsi
    pip install -e “.[dev]”
    python3 rrsi.py –domain coding baseline
    python3 rrsi.py –domain coding run

    Each round drafts 2 candidates in separate git worktrees, screens them, evaluates both and fast-forwards the branch to the winner. The coding instance also needs Docker and harbor. New domains plug in through a single adapter module.

    Key Takeaways

    • RRSI evolves prompts, tools, memory and workflows while model weights stay frozen.
    • A leakage critic, noise floor, cost rule and pruning decide which edits stick.
    • Terminal-Bench 2.1 rose from 74.2% to 80.2% with Claude Opus 4.8.
    • SWE-bench Verified, never used for selection, rose from 82.0% to 83.8%.
    • Apache 2.0 code on GitHub; research-grade, not an official Google product.

    Check out the Paper, Codes and Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

    Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.



    Source link

    Previous ArticleRaydium’s Holders Revenue Reaches $4.43M as StonkFun Drives LaunchLab Volume

    Related Posts

    New formulation helps RNA vaccines withstand high temperatures | MIT News

    September 28, 2026

    Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building

    September 26, 2026

    Estimating suicide risk from text | MIT News

    September 25, 2026

      Subscribe to Updates

      Subscribe to our newsletter for early access to new products, exclusive deals, and exciting updates. Don't miss out! Our subscribers are always the first to hear about limited-time offers and new arrivals. Plus, you'll get sneak peeks and bonus content that adds value to your experience.

      By opting in you agree to receive emails from us and our affiliates. Your information is secure and your privacy is protected.

      Top Posts

      Raydium’s Holders Revenue Reaches $4.43M as StonkFun Drives LaunchLab Volume

      September 29, 2026

      DOGE Rally Brewing? Whales Load Up as Dogecoin ETFs Post Record Inflows

      September 29, 2026

      BCH and NEAR Surge 34% as Altcoins Outpace Bitcoin

      September 28, 2026

      Investhena is a digital news blog covering the latest updates in crypto, global economy, and investing. We focus on clear, timely insights to help readers stay informed and understand market trends without unnecessary complexity.

      Letest News

      Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting

      September 29, 2026

      Raydium’s Holders Revenue Reaches $4.43M as StonkFun Drives LaunchLab Volume

      September 29, 2026
      LEGAL INFORMATION
      • Contact us
      • Terms & Conditions
      • Privacy Policy
      Copyright © 2026 investhena.com | All Rights Reserved

      Type above and press Enter to search. Press Esc to cancel.