birds dancing on my ceiling
13
75
709
Clockspring
2
5
68
1,440
max retweeted
Live now: GPT-6 Astra and Claude 5.5 Opus are racing to create the best StarCraft strategy x.lingyaoai.com/i/broadcasts/1nxeLMvdo…
70
159
1,386
277,676
Want to remove algorithmic feeds from your phone, but still have access to your DMs? Install Beeper, perfect for x + ig messaging without the distraction
2
24
2,448
max retweeted
Introducing Blender Bench v1 (BBv1): a benchmark for frontier LLMs on 20 realistic 3D production workflows including animation, modeling, rigging, cloth & surfacing. GPT-6 Astra leads overall; Claude Opus 5.5 wins modeling & cloth simulation.
47
57
651
40,327
max retweeted
multi agent social strategy is much harder to benchmax because it’s about how the models perform against each other rather than one’s test grading it’s emblematic of real world competition. it’s why swarms and multi agent is lucrative for capability but worrying for safety.
Replying to @sensho
Really promising benchmark, how does it get benchmaxxed?
4
1
30
2,079
The most important messages often don't reach people, The most viral memes are almost meaningless. Like the area of rectangle, the path to the highest impact is improving whichever side you're shortest.
24
1,230
I think Runebench is directionally useful signal because: - research->act->optimize is a realistic loop - It's highly time limited, models need to work quickly - It's not (yet) benchmaxxed to hell
Replying to @maxbittker
this is unironically the benchmark i pay the most attention to now, outside of my own usage
4
1
78
3,890
GPT-6.1 Sol scores #2 on Runescape Bench, at ~10% the price of Astra. Look how far it is above the Pareto efficiency curve and note log-scale Y axis
18
33
561
32,219
On high thinking, GPT-6.1 Sol solved "Waterfall Quest" in both Attack and Strength runs for huge XP drops! (other top models have tried this and failed)
3
27
2,665
Watch GPT-6.1 Sol high completing Waterfall Quest This run was was 30 minutes wall-clock time. (The game run at 8x speed, and the video is another 10x on top) Astra, Opus, Gemini, and Grok have all attempted this quest and failed!
4
465
That ambition cut both ways- here's GPT-6.1 Sol's Crafting run, where it tries to complete elemental workshop and dies + runs out of time.
2
12
1,978
It's great to see labs competing on efficiency These were 30 minute trials, so you can expect that at API prices, new Sol costs roughly $4/hr to Astra's $30 or Opus's $8 Full trajectories videos and stats here: maxbittker.github.io/runeben…
10
1,362
Super excited for this talk! If you've never been to Recurse Center, it's a great opportunity to visit the space
Can automating an MMO teach us about multi-agent systems? On October 14th, @maxbittker will present RuneScape Bench: MMOs for Agents. RS-sdk is a 2026 project to turn an open source implementation of the game, LostCity, into a testbed for new experiments. Its results have implications in many areas: design libraries for coding agents, evals, LLM post-training, game design, simulated economics, and even the future of multi-agent systems, when agents are tasked with trading and trusting one another in a shared world. RSVP below:
7
11
181
8,471
Can automating an MMO teach us about multi-agent systems? On October 14th, @maxbittker will present RuneScape Bench: MMOs for Agents. RS-sdk is a 2026 project to turn an open source implementation of the game, LostCity, into a testbed for new experiments. Its results have implications in many areas: design libraries for coding agents, evals, LLM post-training, game design, simulated economics, and even the future of multi-agent systems, when agents are tasked with trading and trusting one another in a shared world. RSVP below:
4
9
84
12,833
Swarm Scaling Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm? 🧵
25
59
401
73,844
In a multi-agent evaluation on ProgramBench: Sonnet 5.5's fastest setup was subagents Opus 5.5 went fastest using a 5 agent team
2
3
28
2,126
This reminds me of @polynoamial 's hypothesis that working together relies on a base level capabilities in the individual model.
1
6
386
cc @jyangballin @OfirPress so interesting seeing programbench being used this way
1
2
263
figures from the recent sonnet and opus system cards, recolored for clarity.
2
130