yells at @theo both on and off camera (co-host of the nerd snipe podcast)

San Francisco
Pinned Tweet
My vid on Opus 5.5 and GPT 6.1 Sol: - Sol 6.1 is miles better than 6.0, it's worth trying out & the price is nuts - Opus 5.5 is the best Claude model in a very long time
17
3
295
15,743
Ben Davis retweeted
All right, we’re locking in. Only things being worked on are simplifications, more efficiency for more usage, groundbreaking features or new models. Sometimes you have to invest ahead of the curve, but feedback is clear that you all want things to get simpler. On it.
1,688
284
12,371
835,764
I'm enjoying Astra on low reasoning + ultra fast way too much for a ton of little things like: - email management - one off fixes - research - quick demos Usage limits are still pretty brutal. Really want GPT-6.1 Sol ultra fast that's gonna be absurdly useful
72
17
1,261
278,860
Ben Davis retweeted
I deeply respect anyone who can update their views in the face of new evidence. The last four years have really taught me how rare it is for someone to be able to publicly change their mind.
Claude Ultracode is super intelligence.
132
259
5,646
283,392
Both SvelteKit v3 and Effect v4 just released They're excellent releases. I've been gutting my stack over the last few months and effect + sveltekit have become two core pillars of basically everything I work on Highly rec throwing ur agents at these to try them out...
10
9
305
14,645
Ben Davis retweeted
SvelteKit 3 is out! Better environment variable handling, better error messages and sensible breaking changes with an easy upgrade thanks to our migration script which handles many things automatically and gives you(r agent) a nice Todo list for the rest svelte.dev/blog/sveltekit-3-…
18
77
417
22,537
This one simple trick makes ultra fast ultra fast:
8
271
10,353
The approvals really slow it down. 99% of the time auto mode is correct, just like 99% of the time ultra fast is dumb But for that 1%...
24
1,703
The only real option in the "personal agent" space right now is Grok Bot. Been doing a lot with Dots. I do not like them. Dots are a total mess right now. 1) Having a single thread for everything sucks. They got it right first try with Grok Bot, I have zero clue why they wanted it to just be one thread (I do know why, it's because I'm not the target market for Dots but still I hate it) 2) It's slow. Very, very slow. 3) There's zero visibility into what it's actually doing. All u can see is the final short text message that Astra gives u (5-10 mins after asking something). Grok bot's pretty good about talking u through what it's doing and being pretty fast, Bot just isn't 4) The app(s) are a mess. The mobile app, the desktop app, they all have performance issues, weird ux things (like when ur not focusing the input box on desktop and type it doesn't auto focus), just an absurd number of paper cuts. It's weird b/c Cursor is just as bad about this kind of thing, but Grok Bot is absurdly polished 5) I have zero clue wtf this thing is doing. I woke up to a bunch of random codex threads doing god knows what in circles and I have zero trust for this thing if I'm gonna be honest. 6) Zero parallelization. No ability to run a bunch of stuff in parallel, take advantage of the thing Astra's unbelievable at (agent swarming), just one little thread with 5 subagents. I get that I'm not who this is built for. This isn't built for professionals or technical people. It's meant for someone to just chat with an AI and have that AI learn about them and do little useful things. The only one of these that you can do actual serious work on is Grok Bot right now. If you are employed or do anything useful, use Grok Bot. My read on the space right now is that Muse will dominate the bottom 90% b/c free IG distribution and it doesn't need to be good to do what 99% of people will use it for. Dot is in this bizarre middle ground rn where they want it to be something accessible to everyone, but also good for serious work. The "middle" in product is always fucking stupid and doomed. Grok Bot is the "serious" option, just use it. I keep saying this but SpaceXAI is the best engineering lab right now and it's not close. Compare grok dot com, their mobile app, grok bot, and grok build to their OpenAI or Ant equivalent. It isn't close, but then they release Grok 4.7 and it's a useless, overpriced, dumpster fire. I hope Dot improves, they've got the models to really make it happen. But at least for right now, it's a mess and not something I would recommend.
123
31
765
42,169
Effect v4 is here. One ecosystem. Zero dependencies. The next chapter of Effect and the foundation for building reliable software and AI agents in TypeScript.
79
320
2,309
369,683
Sickness has finally gone away and I'm functional again, just recorded my GPT-6.1 vs Opus 5.5 vid They're both great, both feel nearly unlimited, and my favorite recent example of OpenAI/Ant fighting resulting in us getting better products Will be out tmrw morning 🫡
8
2
249
31,075
Ran the numbers on Sol v Opus for "cost per task" (wouldn't take "task" too seriously it's just whatever Sol thought was a task) The price cut on 6.1 Sol is kinda insane, especially when u combine it with how efficient OpenAI models are Opus still has that special "something" that good Claude models are known for and in the $200 Claude plan has great usage limits But for a model that's 90% as good, doesn't have any obvious major issues, and is effectively unlimited 6.1 Sol is a very slept on release
1
6
2,669
Ran the numbers on Sol v Opus for "cost per task" (wouldn't take "task" too seriously it's just whatever Sol thought was a task) The price cut on 6.1 Sol is kinda insane, especially when u combine it with how efficient OpenAI models are Opus still has that special "something" that good Claude models are known for and in the $200 Claude plan has great usage limits But for a model that's 90% as good, doesn't have any obvious major issues, and is effectively unlimited 6.1 Sol is a very slept on release
Sickness has finally gone away and I'm functional again, just recorded my GPT-6.1 vs Opus 5.5 vid They're both great, both feel nearly unlimited, and my favorite recent example of OpenAI/Ant fighting resulting in us getting better products Will be out tmrw morning 🫡
31
5
348
27,454
I did have to reset my limits today on Codex, but that's just because of a certain ultrafast which will incinerate your limits like you wouldn't believe
23
1,853
You fools mocked the Googlebook, but I have always believed
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
25
1
296
12,883
I really hope it's actually good it would be fucking funny
5
67
1,702
Had Astra on ultrafast (on med reasoning) rebuild 1-1 from scratch on the modretro they gave us at Dev Day yesterday - Took ~3 mins - Used 4 subagents - Used ~10% of my weekly limit ($500/mo. plan) Ultrafast should rarely be used, but I'm ngl when u do use it it's really cool
27
10
366
39,617
In actual work, I think I'll probably gonna end up using ultrafast for when 1) I really want to not be context switching and 2) it's the type of thing that can/should be done on light reasoning (email, notion, report building, etc.) When u have 0 sub agents and lower reasoning levels, it's less brutal on ur usage limit than I was expecting (but still very noticeable)
3
18
1,678
Ben Davis retweeted
If you’re a Figma user trying to get work done but can’t because of their shitty MCP setup, it’s worth considering switching to @paper who don’t have the same restrictions
today in MCP land ... thing are better compared to a year ago, but also worse.
32
6
357
25,134
Ben Davis retweeted
Usually demos like these have someone back stage doing the exact same thing, ready for instant cut over. If that also fails, a pre-recorded video is the 2nd cut over. This made me like OpenAI even more. Live demos are so hard, the internet in this venue can be flakey, you are super nervous, etc.. This stuff takes weeks to prepare and I doubt these presenters had that much time
HOLY SHIT OpenAI's live demo for "Dots" completely FLOPPED💀 and then they CUT THE AUDIO on the official stream to hide it I caught the original:
52
13
879
115,264