Devices @OpenAI

San Francisco, CA
Peter Welinder retweeted
gpt-6.1 sol has saturated the FrontierMath Tier 4 benchmark that was considered exceptionally difficult even for research-level math just last year that's only 10 months of progress also Astra solved 5 erdős problems, and 2 of them under budget
34
63
819
58,352
Peter Welinder retweeted
Who the fuck comes up with this shit lmao
Mom
63
262
3,125
327,093
Peter Welinder retweeted
Just this morning, needles and other drug paraphernalia reported next to Marshall Elementary. "A couple weeks ago, a first grader reached through the fence surrounding the recess yard and grabbed a bloody hypothermic needle. He passed it around to four other children before it was caught."
Parents and staff of Marshall Elementary in the Mission District and the area’s supervisor said street conditions have worsened since school started. They’re asking for more help. sfchronicle.com/sf/article/s…
7
20
68
13,242
Peter Welinder retweeted
wowi GPT-6.1 Sol scores 100% at FrontierMeth Tier 4 v2; shoutout to @AcerFur xoxo
19
15
432
20,984
Astra reveals previously unknown history
I used GPT-6 Astra to break an unsolved cipher to one of Napoleon's generals that had gone unread for 217 years. What makes this impressive isn't actually the codebreaking, but that Astra completed the entire multi-modal workflow in ~6 hours from a single image and goal. 1/6🧵
8
9
158
50,370
Peter Welinder retweeted
I’ve been running a bunch of tests with GPT-6.1 Sol, Astra, Blender, and Three.js, and after a few hours of iterating, the results have been kind of insane. There’s really no denying that 6.1 Sol is excellent, especially when you give it the right kinds of tasks. OpenAI has never cooked this hard with a model before. (And that’s coming from someone who consistently posts criticism when they get things wrong.)
27
8
183
11,542
Peter Welinder retweeted
GPT-6.1 Sol is the new #1 on MathArena! Cheaper and better than Astra.
21
52
520
27,613
Peter Welinder retweeted
GPT 6.1 Sol in Codex continues to be the absolute goat for mobile app development
19
12
172
12,894
Peter Welinder retweeted
Dots are here! A new way to use AI that works 24/7 for you; get more of your time and attention back to work at a higher level. openai.com/index/introducing…
2,122
1,081
14,752
2,026,510
Peter Welinder retweeted
Ultrafast - 8x the speed, up to 300 tokens/second. once you go ultrafast you won't go back :)
21
5
92
94,535
Peter Welinder retweeted
1 year after the original codex launch, codex cloud is back!
Live from the Codex Cloud Warroom: we are at 100% rollout for pro users!
6
4
72
11,674
Peter Welinder retweeted
🎯
openai announces glm-5.3 in codex subscriptions anthropic announces glm-5.3 is evil and chinese
1
4
94
26,274
Open as in open source. We just want to enable you to build.
Enterprise teams can now use open models like GLM-5.3 Flash and Kimi K3 natively in Codex and count spend against their OpenAI commit.
7
2
41
6,254
Personal stylist in ChatGPT
I built a lil personal stylist for ChatGPT. Mens and womenswear. Upload your photo and get 14 complete looks with shopping links that fit your budget. In just 4 mins. Submitted for review. You can play once approved.
2
6
21
5,360
Peter Welinder retweeted
GPT-6.1 Sol scores #2 on Runescape Bench, at ~10% the price of Astra. Look how far it is above the Pareto efficiency curve and note log-scale Y axis
18
33
561
32,239
Much better than #3 and 1/8th the price 🔥
GPT-6.1-Sol gets #2 on eyebench-v3, stealing Opus-5.5s short-lived spot! It's near-astra performance for ~3.8x cheaper, and ~8x cheaper than Opus-5.5 while scoring better.
1
2
36
4,678
Peter Welinder retweeted
we found opus 5.5 to be 2X as efficient as gpt-6-sol in our internal cfo.ai evals so i was not expecting 6.1-sol to improve on that much but to see it 3X the efficiency of opus 5.5 and 2.5X against even sonnet 5.5 is actually 🤯 this is an insane point release
133
244
4,279
326,148
Peter Welinder retweeted
So I tested GPT-6.1 Sol on real work: 2 repos, 105 planted bugs. Find and fix what you can. Unlike GPT-6 Sol, which was just a nerfed GPT-5.6 Terra, this one is real. The results: - GPT-6 Astra (max): 45 for $33 - GPT-6.1 Sol (max): 44 for $6.56 - GPT-5.6 Sol (max): 43.5 for $95.35 - Opus 5.5 (max): 41.7 for $58.53 - GPT-6 Sol (max): 29.3 for $9.33 With a model like this, the 50% cut to the $200 plan doesn't matter. n=1. More runs and results at more effort levels dropping in this thread over the next few hours 🧵
232
207
3,107
407,743
Peter Welinder retweeted
GPT-6.1 Sol ULTRA ran for 25 minutes and only used 1% of my weekly quota, It’s seriously fast and powerful, almost like Astra, but much cheaper to run, For me, this is the kind of model you can actually use 24/7 without constantly worrying about limits, From now on, it’s probably going to be my default model, Honestly, out of everything announced at DevDay, this might be my favorite update. In my experience, it feels like the most efficient model OpenAI has released so far
72
40
1,005
140,842
Figma is awesome. Now available in Codex.
Dev Mode is now available in Codex! You can see what’s ready for dev with the Figma plugin and compare to local builds with this new MCP functionality.
1
11
2,751