AutoResearch is critical in unlocking new knowledge and accelerate discovery. And it's an important ingredient in our quest to RSI.
Today we at
@bespokelabsai are happy to announce a new benchmark that's tailored to measure agents' ability to do autoresearch: AutoResearchExam.
As part of this benchmark, we release 29 tasks that measure progress over 24 hours for agents to do sustained ML research and model training.
Please check out for more info:
benchmarks.bespokelabs.ai/au…
This line of work is especially timely, given the recent advances in math and science, such as solving the Navier-Stokes Millennium Prize Problem, which needs sustained autoresearch and test-time scaling.