The official AI benchmark of the vibe coding movement @bridgemindai

United States
Bridgebench retweeted
@theo, it's very clear you're jealous of BridgeMind. Some of your questions were fair. I'll be sharing more on the NerfBench methodology soon. But asking me to take NerfBench down and post an apology written by you isn't about better benchmarks. It's about your ego, and not wanting anyone else to exist in this space. More benches are good for everyone. Build yours.
Replying to @theo
I'm down to put my money where my mouth is. BridgeMind put $10,000 towards benching, I'll do the same. If my benches show I'm wrong, and that there is meaningful model "nerfing" aligned with his viral post, I will donate another $10,000 to a charity of @bridgemindai's choice. If my benches show I'm right, I want Bridgemind to take down "nerf bench" and replace it with a page apologizing, listing all the reasons the bench was flawed (written by me). I'm even down to have a qualified third party come in and audit both of our work. Deal?
178
31
1,376
122,892
Bridgebench retweeted
The OpenAI situation keeps getting worse. Yesterday Tibo announced a global reset for all paid ChatGPT accounts at 10am PST today. It's been over 2 hours. No reset. Builders literally scheduled their whole day around this. First they cut usage in half. Now they promise a reset that doesn't show up. You can't keep doing this to the people paying you.
256
44
1,424
94,389
GPT 6.1 Sol is 2.5x slower than GPT 6 Sol. Same lava lamp prompt: GPT 6 Sol: $0.08, 50s GPT 6.1 Sol: $0.09, 2m 9s OpenAI marketed 6.1 Sol as more efficient. On BridgeBench it took longer and cost more.
74
16
604
60,395
Bridgebench retweeted
Claude Opus 5.5 just one shot a full multiplayer battle royale like Call of Duty Warzone. Then I used Claude Opus 5.5 with Ultracode to polish it. Parachute drops, a storm that closes over 8 circles, a full map with named towns, and real players dropping in from around the world. Opus 5.5 is unstoppable.
130
68
1,516
97,881
Anthropic could nerf Opus 5.5 to 80% power and it would still be way better than GPT 6.1 Sol. OpenAI is that far behind.
66
14
617
19,660
Bridgebench retweeted
Claude Opus 5.5 just took a big drop on NerfBench. Yesterday it was scoring above launch. Today it's at 94.2%. GPT 6 Astra: 98.0% Sonnet 5.5: 100.9% GPT 6.1 Sol: 106.7% 94.2% is still inside normal variance, so we can't call it a nerf yet. But we're watching Opus 5.5 very closely.
377
332
6,759
760,780
Bridgebench retweeted
Fable 5.5 is coming and it's going to be one of the biggest jumps in capability we've ever seen. Opus 5.5 already matches Fable 5.1 at 40% of the price. Imagine what Fable 5.5 does. But it has to be affordable. Fable 5.1 is $10 in and $50 out and Claude Max dies in 30 minutes on it. That's not usable. Fable 5.5 has to be good. It also has to be something we can actually use.
34
11
401
27,573
Claude Opus 5.5 is the best front end design model on BridgeBench. Anthropic holds the top 4 spots. OpenAI doesn't have a single model in the top 10. Not GPT 6 Astra. Not GPT 6 Sol. Not even the new GPT 6.1 Sol. Design has been OpenAI's weak spot for 3 years and they still haven't fixed it. When will OpenAI finally fix design?
52
9
414
21,102
Bridgebench retweeted
Fable 5.5 is going to be the biggest leap we have ever seen. Anthropic's IPO is coming in November. They are going to build maximum momentum before it, and Fable 5.5 is how they do it. OpenAI is not ready.
143
93
3,355
124,077
Opus 5.5 is leagues ahead of GPT 6.1 Sol. When Anthropic drops Fable 5.5 OpenAI is cooked.
61
17
990
34,258
Bridgebench retweeted
One GPT 6 Astra agent on Ultrafast drained my entire $500 plan for the week in 15 minutes. Then it burned through 95% of the $2,500 in credits OpenAI gave me for cutting my usage in half. 3,103 credits left out of 62,500. OpenAI is terrible.
189
53
1,507
80,003
Claude Opus 5.5 and GPT 6 Astra are not nerfed. Latest NerfBench results: Claude Opus 5.5: 103.8% of launch power GPT-6 Astra: 101.1% of launch power Both are within normal variance (90% to 110%). We also just added GPT 6.1 Sol. Day one baseline is locked in. Now we watch. Which model should we add next?
123
64
1,825
64,080
Bridgebench retweeted
Google has been quiet for months. Now this. Gemini 4 Argon has a 15% hallucination rate. The lowest on the entire board. Opus 5.5: 59%. GPT 6 Astra: 51%. Fable 5.1: 73%. And Gemini 4 Argon is not even publicly available yet. How is this possible?
166
60
1,956
135,014
Bridgebench retweeted
I'm genuinely really happy with Anthropic right now. OpenAI just cut usage by 50%, and I'm really glad we have Opus 5.5. It's the best model I've ever used. It's fast, it's $4 in and $20 out, and I'm not worrying about limits anymore. It's made me about 2x more productive. I'm optimistic Anthropic keeps it this way. For the first time in a while it feels like anything is possible.
34
18
583
26,458
GPT 6.1 Sol vs Claude Opus 5.5 on the BridgeBench turntable test. Opus 5.5 built a full 3D turntable. Wood grain, vinyl, tonearm, all of it. GPT 6.1 Sol didn't produce a usable result. Near GPT 6 Astra on OpenAI's benchmarks. A blank screen on ours. Anthropic is winning right now.
61
10
452
24,475
Bridgebench retweeted
I JUST CANCELLED ALL OF MY CHATGPT SUBSCRIPTIONS. OpenAI just scammed everybody and BridgeMind is now starting a boycott against OpenAI. They cut the usage in half yesterday and said better efficiency makes up for it. This is a lie. GPT 6.1 Sol is just GPT 6 Sol with RL to look better on benchmarks. If we keep paying, Anthropic and every other lab will follow. Cutting usage is profitable for them. Do not let OpenAI get away with this.
589
362
4,870
393,865
GPT 6.1 Sol is now on BridgeBench. OpenAI is calling it "near-Astra intelligence." Our results: 3 points better than GPT 6 Sol. 82 points behind GPT 6 Astra. When a model jumps on the lab's own benchmarks but barely moves on independent ones, that's usually a sign of benchmaxing.
62
56
1,347
85,771
OpenAI is marketing GPT 6.1 Sol as "near-Astra intelligence for a fifth of the price." We ran it through the BridgeBench sunset ocean test. Same cost as GPT 6 Sol. Twice as slow. And it barely rendered an ocean. This isn't Astra level. It feels like GPT 6 Sol with some better RL.
96
41
1,207
153,606
Bridgebench retweeted
One GPT 6 Astra prompt using ultrafast on the new $500 ChatGPT Pro Max plan. Ran for 2 minutes. Used 2% of my weekly limit. That is 50 prompts, or about 100 minutes of work, before the week is gone. For $500 a month. This is a complete scam.
134
113
2,026
95,085
Bridgebench retweeted
I am shocked that OpenAI would rug pull subscriptions by 50% right before OpenAI DevDay.
64
37
1,214
39,737