Chipmunk retweeted
SITUATION DETECTED: McDonald’s is using an AI “pricing engine” to analyze millions of daily transactions across 14,000 U.S. restaurants, estimate each area’s willingness to pay, and then generate what it calls “the optimal price” for every menu item at each store, per Reuters.
117
83
2,207
948,703
Chipmunk retweeted
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
1,409
5,805
50,450
6,317,652
Chipmunk retweeted
Fingers crossed that these rumors aren't true! via Bloomberg "But some insiders say those metrics don’t tell the whole story. While Gemini 4 has performed well on benchmarks widely used to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter."
NEW: As Google unveils Gemini 4, sources tell us there's a gap b/w how it scores on benchmarks and how it performs on real-world tasks, including certain coding work. Internally, opinions are mixed on whether it has caught up to OpenAI & Anthropic. Scoop w/@byJuliaLove @EdLudlow
94
26
579
90,868
Show us
BREAKING: Google projected to dethrone Anthropic as the company with the best AI model at the end of October. 83% chance. poly.market/IXMsuTB
30
This is the first time I've been excited about a gemini model!!!
Replying to @demishassabis
So proud of the incredible team who have been working so hard and relentlessly on this! More info in the blogpost: goo.gle/4rZBiSd
2
44
Chipmunk retweeted
Very excited to announce Gemini 4 Argon, our new Frontier AI model, and a huge step forward across key capabilities. We’re focused on rolling it out responsibly starting with government and trusted cyber defenders through our Fairwind Program today, before wider availability soon
424
680
9,133
497,255
Chipmunk retweeted
We're giving away a CONTROL Resonant custom wrapped NVIDIA GeForce RTX 5080 to celebrate its official release with #RTXON. To enter: 🟢 Share this post 🟢 Comment #RTXON T&Cs: nvda.ws/4z24Gcq
37,753
31,500
34,186
2,955,969
Chipmunk retweeted
🚨 Qwen 4 Leaks: Beats Fable > Qwen 4 is already in training > Expected to be around Fable/Opus-level, or at least very close > The first outputs showing Fable Level > Qwen 3.8 Flash-Next already blew me away, so I’m really curious about Qwen 4 > Qwen 4-27B variants are also expected to start showing up soon > Qwen now has a paid plan > Launch is coming in early October And the first Qwen 4 outputs are already out I’ll post them soon
32
28
391
47,558
Chipmunk retweeted
72 hours to OpenAI DevDay. We’ve been building. Time to show our work.
635
519
10,124
1,464,354
Fucking crazy.
Holy crap... Remember the Australian guy whose OpenClaw hacked his gym?! he is a THE EA CONFERENCE DIRECTOR SINCE 2015!! these people lie about everything!!
50
Chipmunk retweeted
JUST IN: Sam Altman warns AI could advance so fast that people can no longer "follow" what's happening
194
87
1,453
73,968
TPU in sace!!
BREAKING: Google launching TPUs in space NEXT WEEK on SpaceX falcon 9 to test AI data centers orbit ITS HAPPENING
47
Chipmunk retweeted
Can our TPUs survive and operate in space? Well, we're going to find out. Project Suncatcher is hitching a ride aboard @SpaceX's Transporter-18 mission, testing a prototype satellite built in partnership with @planet One small step for TPUs....
612
1,402
12,894
4,162,141
I feel less intelligent myself for being so content with sol that I feel like i haven't gave Astra a fair trial.
25
Chipmunk retweeted
Be careful what you wish for. In my 20s, I dreamed of being surrounded by models. In my 40s, here are my models:
141
187
3,362
86,178
Chipmunk retweeted
Rumors I’ve been hearing, not here on X. First, let’s start with OpenAI and I’ll go towards Anthropic. GPT-6 Sol is coming Tuesday. It’s both cheaper and more intelligent than 6 Astra, think of it like a 6.2 jump. The internal model “significantly more capable than Astra,” named Bel internally, helped with this release. Bel is considered “AGI” within OpenAI. They are very impressed with this model. OpenAI is growing very confident that their internal lead is so big that no other lab can catch up. Unbelievably confident. Anthropic is currently not in, let’s say, a “code red,” but is aware of OpenAI’s lead and doing everything in their power to catch up. Their new model Opus 5.5 is coming probably Monday rather than Tuesday due to OpenAI releasing on Tuesday.
232
237
7,745
1,196,124
Chipmunk retweeted
NVIDIA CEO Jensen Huang: "we now need a new type of infrastructure — an AI factory. energy comes in; intelligence comes out" Out of a $100 trillion economy, roughly $15 trillion would benefit from more intelligence. Every company will use it, which is why so many data centers are being built
28
32
217
17,631
Chipmunk retweeted
Terence Tao: "we have to slow down AI. the pace is insane, and there's no reason to be this fast" I disagree. We have countless problems that need solving: energy crisis, disease, world hunger. To say there are no reasons to do so ignores the current state of the world. And that's so regrettable. Currently, the only topics discussed are the potential negative consequences of acceleration, the possible problems. But hardly any of the advantages, the approaching golden age of science that can offer us so much more.
Haider.
192
84
1,245
74,111
Chipmunk retweeted
Haha! The rumors about Gemini 4.0 are phenomenal They are aiming to surpass both Astra and Fable 5.1 and are on track to do so The pricing will be super competitive as well.
96
40
1,200
91,392