Everyone cites Voyager as the canonical "self-improving agent." It isn't one. Getting the distinction right matters, because it's the difference between a moat and a pile.
【N1】Voyager (NVIDIA/Caltech, 2023) plays Minecraft by writing JavaScript. GPT-4 proposes its own next task, writes a function to execute it, repairs that function against runtime errors, and a second GPT-4 instance verifies success before the function is committed to a vector-indexed skill library. New functions call old ones. It was the only method to reach diamond tools unassisted.
【N2】Swap GPT-4 for GPT-3.5 and unique-item discovery drops 5.7x. Same architecture, same loop.
【J1】That number is the whole story. The system moves capability, it doesn't create it. Voyager widens itself; it cannot heighten itself. Real RSI modifies the machinery that produces improvement. Voyager's weights, curriculum prompts, verification criteria and primitive APIs are all frozen and sit outside the loop. Accumulation is not recursion.
【J2】And for nearly every agent product shipping today, accumulation is enough. Human brains haven't improved in 10,000 years; everything we have came from externalized, compounding assets. Better still: the ceiling gets raised for you, free, every few months. Building RSI means paying to do what the labs already do, and trading away auditability and rollback to get it.
【N3】But Voyager's own discovery curve flattens late. Retrieval stays top-5 while the library grows, so the usable fraction shrinks. Nothing refactors, abstracts, or deprecates.
【J3】So undifferentiated accumulation is not a moat either. Past some volume, a library that only grows starts losing value. Skills rot, weak primitives get depended on, and the verifier never gets better.
【J4】The defensible layer was never the accumulation. It's whatever organizes and validates it. Making that layer self-improving is the cheap, safe 80% of RSI that almost nobody is building.