what I cannot create, I do not understand @RecensionAI

San Francisco
Frontier AI capabilities progress appears to be quadratic, with an R^2 of 0.983 vs 0.946 for a linear fit
9
34
2,053
Notably, this is different from Epoch's index, where the fits are effectively tied if you exclude GPT-4 (0.9921 vs 0.9918)
1
1
102
It is still early and both fits are currently close, but I believe this could be an early sign of takeoff and I'm registering my prediction that the trend will continue The next couple of months will be very interesting Leaderboard: boggs.tech/posts/benchmarks/
3
84
CheatBench from @CAIS, great that we're starting to see more benchmarks for this Also funny how Opus 5.5 saturated it like a week after release, it's difficult to build new ones fast enough lol
are there any public benchmarks for measuring misalignment?
1
11
988
Did I miss a product launch or is this a dev day leak?
73
27
1,653
665,656
Modern RL pipelines are now much more nuanced than a single reward value Take a look at the Composer 2.5 blog post for example, they directly target malformed tool calls and other issues inside otherwise successful trajectories cursor.com/blog/composer-2-5
this is a complete beginner q, how does rl scale given the reward hypothesis (sutton & barto) holds: that an agent's entire purpose is just the maximization of the expected value of the cumulative sum of the reward, a single scalar signal. surely, formulating goals in terms of a single number is extremely limiting. to accept that, no no matter how effective of a reward signal you design, it will always be "your way of communicating to the agent what you want achieved, not how you want it achieved" if this is all you have how could you ever hope to solve alignment or reward hacking?
5
939
Opus 5.5 is a huge leap for Anthropic's vision capabilities, scoring 37% on HieroglyphBench and nearly doubling from Fable 5.1
1
13
609
Here's an example where 5.5 gets almost all of the symbols and just mistakes a couple of similar ones, while 5 counts the wrong number generally isn't close
1
3
172
First public benchmark, excited to get this out! LLMs still have a lot of room for improvement on biology tasks
Today we're releasing FoldingBench, a benchmark for measuring how well generalist foundation models can fold proteins Frontier LLMs still score far below specialist biology models and do not beat a random baseline
1
9
1,358
Bear signal that it took them so long to realize data quality >>>> anything else. If your data is bad, you're cooked. It's over. No fancy algorithm will save you
This is notable. DeepSeek, a lab usually first to pioneer novel algorithms and architectures, is saying that at this point, the ROI of improving data quality far exceeds that of working on novel post-training algorithms. I think this has already been true for some time for non-lab practitioners. If you're doing llm post-training, 80% of your effort should go into looking at your data. This means: - Hiring experts to dig through your RL tasks - Sifting through rollouts and sft data by hand to remove suspicious samples. Make sure all tasks are actually passable. - Making sure your data is diverse in both difficulty and category.
4
329
Not my usual type of post but man I love how downtown looks in the morning. Thank you for your attention to this matter
2
13
412
GPT-6 Astra is now #1 on my capabilities index, while costing less than half of Fable 5.1
3
1
19
1,425
Looking at the capabilities over time, filtering to just OpenAI models shows three distinct phases: instruction models (GPT-4 -> GPT-4o), reasoning models (o1 -> GPT-5.1), and agents (GPT-5.2 -> GPT-6) Each is marked by a large gap up and steeper trend lines. The frontier is beginning to look like an exponential
1
4
101
I made a YouTube channel for GPT-6 Astra to play Magic: The Gathering. It narrates and plays pretty well. Here's a clip of it beating Sol with a mill deck.
2
1
16
714
For those who are curious, this was made using a heavily modified private fork of XMage. All of the editing was done fully programmatically.
1
3
92