This looks like a cool new coding bench! Basically give the agent a repo at one commit in the past, say "find and fix all bugs" and test against real bugfixes from future commits; see if it found all of them (via their unit tests).
I don't think it's perfect: the model might find 8 bugs but the test is about 8 *different* bugs, then it would score zero there but actually be just as useful as another model which finds only exactly the 8 tested bugs. Also, now that the construction is known, the recipe to start training on test is also kinda clear. But the nice thing is that if model providers do this more broadly, it will still be a useful improvement mk of the model.
But: perfect is the enemy of good, and this is clearly a very useful new swe bench for the near future!
Just, once models get to the high percent scoring range, we shouldn't obsess over small takeovers/ranking diffs.
(And make sure the future history is pruned in the eval envs lol i didn't check, but i think they learned the lesson🤞)
Can LMs discover & fix bugs if you don't tell them what went wrong? To proactively maintain a repo, agents need to find problems before users do.
SWE-sweep benchmarks this on 100 repos, 22 languages, 4k real bugs (numpy, php interpreter, lean kernel). Top models get <5%.