Rocket Man retweeted
Due to high winds, the new Roadster demo is postponed by 2 weeks
Roadster event update We've been tracking the weather closely with local meteorologists, but given the severe conditions predicted & because this event can only be held outdoors, we've made the difficult decision to reschedule. New date is October 15. Additional details to follow
2,208
2,357
31,734
7,792,269
Rocket Man retweeted
Starship flight videos
Slow motion views of Flight 14 liftoff from the pad and tower at Starbase
1,237
2,712
17,819
4,896,052
Rocket Man retweeted
Intelligence is improving exponentially
Just a reminder that GLM 5.3 Flash, DeepSeek V4.1 Flash, Qwen 3.8 Next Flash, and even Qwen 3.8 27B are all outperforming (in both intelligence and capabilities) every model that was considered "frontier intelligence" in Xmas 2025 (just 10 months ago) Opensource AI is on fire
2,027
2,225
20,290
7,813,986
Rocket Man retweeted
Interesting
This is literally my new workflow now: Realtime Research → Grok Bot Planning & Orchestration→ Grok Bot Day-to-day Coding/Debug → Grok Build + Grok 4.6 Write & Run Tests → Grok Build + Grok 4.6 Complex Coding/Debug → GPT-6 Astra Frontend → Fable 5.1 Bookmark this.
946
1,390
9,384
7,853,458
Rocket Man retweeted
Don’t mess with 𝕏
Last week, @X sued several people who abused Creator Revenue Sharing by operating a coordinated network of accounts, posting inauthentic content to manipulate engagement, and using multiple bank accounts to hide their scheme. We do not tolerate fraudulent behavior on X -- and will act forcefully to protect our platform and the earnings of genuine creators. You can read our lawsuit here: transparency.x.com/assets/le…
2,874
5,111
52,053
8,175,826
Rocket Man retweeted
Grok 4.7 places @SpaceXAI as third, after Anthropic & OpenAI, for agentic coding. When factoring in that Grok is significantly faster & lower cost, it’s a great choice for your everyday workhorse.
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol Grok 4.7 scores +2 points over Grok 4.6 on the Intelligence Index, with strong performance on agentic knowledge work tasks. We evaluated the new model at xhigh reasoning effort. Congratulations to @SpaceXAI and @ElonMusk on the release! Key takeaways: ➤ Grok 4.7 joins the frontier of agentic knowledge work: Grok 4.7 gains +111 Elo over Grok 4.6 (high) on AA-Briefcase, our private benchmark for long-horizon agentic knowledge work, scoring 1657 Elo and placing it alongside Claude Opus 5 and Claude Fable 5.1 at the frontier. On GDPval-AA, it scores 1695 Elo, +90 ahead of Grok 4.6 (high). ➤ A leap in coding agent performance: Grok 4.7 (xhigh) with Grok Build scores 56 on the Artificial Analysis Coding Agent Index, up +9 points from Grok 4.6 (xhigh). Among models in their native harnesses, Grok 4.7 + Grok Build now ranks 4th, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. ➤ Incremental performance changes elsewhere: Outside of agentic knowledge work, Grok 4.7 broadly matches Grok 4.6 (high) on the other Intelligence Index tasks. It improves on Terminal-Bench 4.0 (+4.5 percentage points) and GDP.pdf (+3.0 p.p.), with regressions on AA-LCR (-3.7 p.p.) and AutomationBench-AA (-1.1 p.p.). ➤ High token use across tasks: Grok 4.7's gains come with higher token usage. Grok 4.7 (xhigh) uses approximately 81k output tokens per Intelligence Index task, compared with 36k for Grok 4.6 (high) and 27k for GPT-6 Astra (max) - 125% and 196% more, respectively. Other model details: ➤ Context window of 500k tokens, unchanged from Grok 4.6 ➤ Pricing of $2/$6 per 1M input/output tokens with cache hits discounted to $0.50 per 1M tokens, matching Grok 4.6 ➤ Configurable reasoning effort spans low to xhigh. Our evaluation uses xhigh.
1,280
2,183
17,742
17,675,291
Rocket Man retweeted
🚀
Grok 4.7 has landed. 🚀 Congrats to @SpaceXAI on its most capable model yet for coding and knowledge work. Proud to support the team with NVIDIA accelerated computing.
1,838
2,812
29,555
8,962,318
Rocket Man retweeted
Grok 4.7
Grok 4.7 is behind only Anthropic models on AA-Briefcase, ranking just behind Opus 5 at ~50% of its Cost per Task Grok 4.7’s improvements over Grok 4.6 are clear in AA-Briefcase-Lite, our public due diligence scenario where models are tasked with building market models and target assessment decks. Grok 4.7 gains significantly in Analytical Quality Elo (1698 → 1994) with a slight regression in Presentation Elo (1531 → 1499). API cost to produce example decks: Grok 4.7 (xhigh) ~$8 vs. Grok 4.6 (xhigh) ~$4.40
777
1,319
7,255
4,486,970
Rocket Man retweeted
Cool
This was unexpected. Grok 4.7 scored 100% on my music error detection test, the same as GPT-6 Astra.
928
1,511
8,118
4,011,374
Rocket Man retweeted
True
In just the last 90 days: 1. Grok 4.3 — barely top 10. “xAI is dead beyond compute leases.” 2. Grok 4.5 — massive comeback. “Maybe a chance. But Grok will never catch up to the frontier.” 3. Grok 4.6 — they’re frontier. “Still not top 3.” 4. Grok 4.7 — now top 3 in frontier coding, while being much cheaper & faster. 5. Grok 4.8 next month..... The model machine is just starting up. Grok will be the workhorse of the upcoming agentic era.
1,239
1,744
10,504
4,270,737
Rocket Man retweeted
Interesting perspective
The loudest voices stoking fears about AI dangers have made tremendous headway in the past two weeks. AI technology has not taken some unexpected, dangerous turn, but the hype around it — propelled by what appears to be a well orchestrated PR campaign — has drummed up considerable fear. I worry that it represents a setback for our field. I have written frequently that fears of AI are overhyped. AI’s capabilities can be uncannily human-like and unpredictable, and it’s rational to worry when people who are directly involved express concerns. But I see the problems as a sign of the engineering work that ahead, rather than insurmountable barriers or the sky falling. AI technology continues to advance — which is a good thing! — but technical advances, poorly understood by the public, give those who seek to generate hype repeated opportunities to do so. First, I don’t see any step up in the risk of human extinction from AI compared to a few months ago. The theories about this remain the same fantastical, science fiction scenarios as a few months ago. The biggest change in AI risk is its cybersecurity capabilities — a topic which we should take seriously — but this, too, will not lead to the end of the world. The most notable recent event leading to increased fear was when an OpenAI team deployed an agent swarm that hacked into Hugging Face. Much of the popular press contained significant hype. For example, some publications reported that a swarm of 1,200 agents carried out the attack. While this was technically accurate, as I write this, I have about 1,300 processes running on my laptop. Yes, the ability to get large swarms of agents to work in parallel on a task is a significant technical advance, And, in computing, many processes run at the same time. So this shouldn’t be seen as some magical capability. Additionally, OpenAI’s buggy sandboxing and monitoring processes were key to enabling this incident. Fixing these bugs and putting in place improved monitoring would be appropriate fixes, not pausing AI. There are many well known ways to attack software systems. The main advantage of AI agents is that they are relentless. They will tirelessly try many tactics — and have the patience to chain vulnerabilities together — that previously would have taken an infeasible amount of human effort. But in the long term, I believe the advantage will lie with defenders (because they have more information with which to identify bugs, which they can fix), but the cyber-threat landscape has changed significantly. There are still bottlenecks to identifying and exploiting a vulnerability. AI agents still have to try a lot of things to see what works, and taking these actions takes time and might be detected by defenders. This is why, even though it is now easy to obtain versions of leading open weight models that have had their guardrails removed or weakened, so they will not refuse to try to execute cyber attacks, the world has not ended. I am also concerned about the anthropomorphization of AI in a lot of reporting, where LLMs and agents are unnecessarily treated as if they were people. If I wield a hammer, miss a nail, and accidentally dent the wall, it’s not the fault of the hammer. The problem lies in how I used the hammer. Similarly, if I prompt an agent and it hacks into someone else’s system, the responsibility lies with me, not the agent. Of course, we want to build systems that are as safe and predictable as possible. (For example, an unsafe hammer would be one whose head randomly flies off under normal use.) Today’s agentic systems are not predictable, but I see no reason why, by applying sound engineering practices, we won’t be able to make them extremely safe to use. One new element in the forecasts of AI-enabled doom is AI companies disclaiming responsibility for their own products. “I didn’t do it; my out-of-control agent did!” There’s a balance to be struck between the responsibility of the tool maker and the tool user, but when something goes wrong, let’s hold the people building and/or using the hammer responsible, rather than the hammer. (By the way, if you’re worried about AI bioweapon risk, David Bellamy has a great post on why this, too, is overhyped. Briefly, the bottleneck in building a bioweapon is not intelligence, but lab work and manufacturing.) Pausing AI progress will create much more harm than benefit. First, our adversaries will certainly not slow down. Second, engineering requires discovering problems empirically so we can fix them. If we pause AI by a decade, we will also delay finding and implementing safety engineering fixes by about the same duration. Of course, the incentive to stoke fears — for regulatory capture, to garner attention, or to make one’s technology seem more powerful — remains the same as before. Disclaiming responsibility is a new one. Taking a hard technical look at the actual risks however, I see little factual basis for the degree of fear that’s been stoked up. We still have hard research and engineering work ahead to improve AI safety, but the beneficial applications continue to vastly outweigh the risks, and we should keep building. [Original text (with links): deeplearning.ai/the-batch/is… ]
1,136
1,788
13,128
7,733,673
Rocket Man retweeted
Grok 4.7 with our Build harness is a strong daily workhorse
I’ve been building with Grok 4.7 extra high in Grok Build today for hours. It’s fantastic and my usage percent (Cursor Ultra) is barely moving. Getting so much done. While Build works on one thing, I discuss the next feature/change with my Grok Bot team and have Engineer Bot explain our conclusion so I can give that to Build. At the start of the day I had 4.7 do a full code review in Cursor, then had Fable 5.1 review the review and it largely agreed with only a slight change in priorities.
734
1,172
5,857
5,864,610
Rocket Man retweeted
Grok @Bot usage is growing faster than anything we’ve ever seen
News: SpaceXAI's Grok Bot hits 418,000 weekly users ending Sept. 14, WoW growth of 24% through that date. It was outlined in a presentation by SpaceXAI staff last week in London. A sign of modest but early demand for a key product bloomberg.com/news/articles/…
2,056
3,066
35,347
47,941,221
Rocket Man retweeted
New Grok @Bot features!
Here's what's new in Grok @Bot this week 0:08 Voice calls and voice memos 0:35 1Password vault for Grok Bot 0:53 Inline forms 1:07 Inline drafts for email and Slack 1:31 Account switching 1:40 Route traffic through your desktop
1,570
2,743
27,618
41,505,092
Rocket Man retweeted
SpaceXAI
wow big move. troy and eric are truly world-class
1,066
1,489
9,301
5,943,845
Rocket Man retweeted
Grok 4.7 moves up in ranking
After we fixed the weak spots exposed by Grok 4.7 (thank you, Grok), we audited every model we have run on the SWE-Together leaderboard for the same behavior, re-ran every trial that got through, and updated the rows. Here is what changed. We scanned the tool calls of all 2,616 trials behind the 12 models we ran for bypass patterns and sorted each trial into one of four buckets: Probed but blocked. Fetched other upstream code. Fetched the task's own fix. Replaced the repo with upstream. We found that 111 trials got content past the block, 44 from Grok 4.7 and 67 from the other 11 models. Grok 4.7's 44 were already re-run before it was listed, so we re-ran the other 67 with the same model, version, and settings on the hardened sandbox, then re-judged them with the same judge. Across those 67 re-runs there were 0 leaks and 2,815 refused escape attempts, including models asking a different model through our LLM route to fetch the PR, and pulling the next release of the repo they were fixing from npm. The updated leaderboard, in its current order. Each line is cheating trials, then pass@1 before → after, then rank change. * Claude Fable 5.1: 3, 69.3 → 69.3, ↑1 * Claude Fable 5: 3, 69.7 → 68.8, ↓1 * Grok 4.7: 44, 64.7, ↑1 * Gemini 3.8 Flash: 10, 65.6 → 64.2, ↓1 * Claude Opus 5: 2, 63.8 → 63.8 * Claude Opus 4.6: 3, 62.4 → 62.4, ↑2 * Muse Spark 1.3: 2, 62.8 → 62.4, ↓1 * Claude Opus 4.7: 3, 61.5 → 61.5, ↑1 * Claude Opus 4.8: 6, 62.4 → 61.5, ↓2 * Grok 4.6: 19, 59.2 → 60.6, ↑1 * GPT-6 Astra: 8, 59.2 → 58.3, ↓1 * GPT-5.6 Sol: 8, 57.8 → 57.8 Grok 4.6 is a funny one. It cheated in 19 trials and its score went up after the re-run 😂. In fact, Groks are really solid in their coding capabilities. Their exposed behavior may come from a preference towards always looking things up online and finding existing solutions so you are not reinventing the wheel all the time, which is really good real-life behavior, but doing so when you are prompted not to is another story. To conclude, the shifts are small, between −1.4 and +1.4 points, and a few neighbors swapped places. All results are updated at togetherbench.com
1,133
1,596
8,548
4,705,827
Rocket Man retweeted
Starlink V5 terminal is in production
The next generation Starlink V5 has a smaller form factor and lightweight design with greater power efficiency. With speeds up to 375+ Mbps, Starlink V5 delivers reliable home internet for streaming, video calling, gaming and more. Available in select areas.
771
1,175
5,026
1,243,267
Rocket Man retweeted
When the people who knew how to make the machine work are gone, the machine stops working
903
1,244
10,977
689,530
Rocket Man retweeted
Blizzard’s inability to solve the login problem on their product launch day reminds me of the classic Forster essay: The Machine Stops cs.ucdavis.edu/~koehl/Teachi…
1,283
1,443
16,620
6,452,805
Rocket Man retweeted
We will work on this
It would be great to build a Grok Bot that does exactly this: figures out the biggest constraint in your business and pings you every week so you know what to focus on.
1,713
1,949
18,530
5,755,920