Progress has been slow over the past couple of years, but recently it has sped up
We view this as a leading indicator of other bio capabilities and will be watching closely to see if this trend continues with future model releases
Today we're releasing FoldingBench, a benchmark for measuring how well generalist foundation models can fold proteins
Frontier LLMs still score far below specialist biology models and do not beat a random baseline
Some responses are blocked by safety filters, but we correct for this using an item response model
This allows us to rank models without unfairly penalizing ones with aggressive classifiers