We officially opened our office in Sydney last week and when we were on the ground, people kept describing Australia as the "lucky country" based on the 1964 book by Donald Horne. It wasn't a compliment. Horne's point was that Australia got rich on what was in the ground, not what was in people's heads. He postulated that Australia's success came from resource extraction, not technological innovation - and this fear was heightened in the age of AI. At Sunrise, Rick Baker pointed out that Australia's instinct in this new AI era is to build data centers - the same old playbook: Build Infrastructure. Sell the commodity. Let someone else capture the software margin. His call to action was a wake up call for ANZ founders to be more ambitious. Well, from our vantage point at @AnthropicAI, founders from APAC are some of the MOST ambitious:
23
27
722
878,894
In Lewis Carroll's Alice in Wonderland, the Red Queen tells Alice "it takes all the running you can do just to stay in the same place." That’s the current AI race in one sentence. A lead measured in benchmarks can vanish in a few months, and the thing you’re measured against keeps improving too. As of a few months ago, SWE-Bench Verified (our previous old gold standard for agentic coding) sits between 93% and 97% across the top 5 labs. At that point the gaps stop meaning much - anyone can route between models for the cost of a config change. So we had to answer an awkward question in public: why build on Claude? 👇
3
1
16
28,037
That’s how we think about Anthropic too: we’re a frontier lab, and also a platform and an apps company. What we mostly do is ship intelligence primitives, and each one, if we’re right, makes a new kind of startup possible. 👉 2024: the API. It made coding agents possible. @cursor_ai, @Lovable, @emergentlabs and @FactoryAI came out of that. 👉 2025: MCP. It let AI connect safely to sensitive systems, which is what regulated, high-trust work in legal, compliance, insurance, and healthcare was waiting for. @harvey, @RogoAI, @OpenEvidence, and @WeAreLegora came out of that. 👉 a month ago: the Model Hardware Standard, an open standard that lets agents operate equipment they’ve never seen before. anthropic.com/news/model-har… Robotics, defense, energy, bio labs. We think it unlocks the next wave of physical AI startups. But it’s early.
2
1
100
If you want to be in these convos, then come join us at Claude Founder House next week during @Techweek_ 👇 🌁 SF: anthropic.com/events/claude-… 🇸🇪 Stockholm: anthropic.com/events/claude-… You can expect: 💡 talks that spark ideas 💡 workshops that turn ideas into plans 💡 a room full of builders pushing each other to aim higher Somewhere right now, a builder is about to change everything - maybe that's you! See you at Claude Founder next week! 👋
82
Jo Zhu Kennedy retweeted
Claude.dev is our new home for developers building with Claude. You'll find engineering deep dives, Claude Code and API guides, tips from the teams building Claude, and some fun easter eggs.
288
829
13,442
1,663,301
Jo Zhu Kennedy retweeted
Claude Founder House is coming to SF Tech Week (Oct 6–8) and Stockholm (Oct 14). Come for talks, workshops, and office hours with our team, plus time to meet others building with Claude. Whether you're a founder, a CEO, or a builder, there's a session for you.
190
176
3,173
281,012
Jo Zhu Kennedy retweeted
SONNET IS ABOVE ASTRA 😭😭😭 what the actual fuck did anthropic do openai you better fucking deliver tomorrow.
107
124
4,459
187,139
Jo Zhu Kennedy retweeted
Sonnet 5.5 fixing a bug with Claude Code. 30% faster and 30% less usage.
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
282
213
5,910
418,258
Jo Zhu Kennedy retweeted
We’re also launching Claude Sonnet 5.5 and Haiku 5.5 in the coming weeks. These models will come with many of the same improvements to performance, efficiency, and safety. With that, happy building everyone.
39
67
1,116
125,752
Jo Zhu Kennedy retweeted
Claude Opus 5.5 is now available in Devin. On FrontierCode 1.1, Opus 5.5 takes the #1 sport from Fable 5 at a fraction of the cost.
44
35
587
80,790
Jo Zhu Kennedy retweeted
Claude Opus 5.5 is now available in Cursor! It's the new top model on CursorBench at 57.8% (Max) and costs 40% less per task than Opus 5.
268
199
4,706
523,188
Jo Zhu Kennedy retweeted
since it’s been a good day for anthropic with a strong opus 5.5 release, i’m going to highlight one more thing that may not be obvious across xai, openai and anthropic - 1. anthropic is the only provider that does not charge 2x for long context requests even at 1M context, anthropic charge at the same flat rate, while openai charges 2x above 272k tokens, and grok charges 2x above 200k 2. anthropic’s latest models have ridiculously low pricing for cached read. see chart below most coding sessions have 95%+ cache hit rate, so this difference is massive my hunch is that eventually this will even out, but for now, this is a very material difference that’s easy to overlook
54
39
754
44,410
Jo Zhu Kennedy retweeted
Today the Pareto curve got updated with new Anthropic Opus 5.5 and OpenAI GPT-6 Sol and Luna models. On the cheaper end GPT-6 Luna is the best for its low end price and Opus 5.5 is the best bang for buck with frontier performance. We are still seeing signs of Fable 5.1 being preferred internally and we are aware benchmarks aren’t perfect, but this sets the new bar for models on the curve It’s also worth noting that other types of tasks are still model dependent as well, we prefer GPT-6 Astra for computer use, as an example, and SWE-2 for model pairings for Devin Fusion since it has been trained to be cost effective across the curve
9
6
78
6,400
Jo Zhu Kennedy retweeted
🚨 Opus 5.5 JUST HIT 58 ON THE ARTIFICIAL ANALYSIS INTELLIGENCE INDEX. The next highest score in this chart is 53. Fable 5.1: 53. GPT-6 Astra: 53. What did Anthropic feed this thing 😭
76
54
1,780
92,268
Jo Zhu Kennedy retweeted
Opus 5.5 is crushing our internal evals
Claude Opus 5.5 is the new #1 on APEX-Agents and APEX-Accounting. APEX-Agents: 73.5% Pass@1 (#1) 81.3% mean score (#1) APEX-Accounting: 15.4% Pass@1 (#1) 62.0% mean score (#1) On APEX-Agents, the new model gains +4.9 pp over Fable 5.1 (68.6%), the previous leader, and +7.6 pp over Opus 5 (65.8%). Anthropic says the biggest gains in this release are on long-running agentic tasks and knowledge work. That is what APEX-Agents measures, and the numbers agree. Opus 5.5 on APEX domains: Management Consulting: 80.0% Pass@1 (#1) Corporate Law: 71.2% Pass@1 (#3) Investment Banking: 69.3% Pass@1 (#2) Accounting: 62.0% mean score (#1) Token use can indicate domains where the model is most effective. Consulting leads at 80% and uses the fewest tokens at 2.0M per attempt. Law has the highest partial credit (85.7% mean) but costs the most at 4.9M tokens. In Accounting, Opus 5.5 gets partial credit on most tasks but fully passes only 1 in 6. More tokens can increase capability on Opus 5.5. Max effort consumes 3.50M tokens per attempt. That’s 2.1x more than Opus 5 and 1.2x Fable 5.1. But list price fell to $4/$20 per M (Opus 5 was $5/$25), and cache reads are $0.20, so 2.1x the tokens is only about 1.7x the dollars. Effort level has a significant impact on benchmark scores and token usage. Medium effort: 52.3% on 756k tokens Max effort: 73.4% on 3.50M tokens Increasing effort to max gains 21 points for 4.6x the tokens on agentic work. On APEX-Accounting, the same jump buys 2.8 points for 3.4x the tokens. When Opus 5.5 solves a task, it solves it consistently. 142 of 239 tasks passed on all 4 runs. Failing runs used 3.1M tokens vs 1.9M for passing runs, and took twice as long. 13 runs hit a hard failure, all in Corporate Law. 8 of them burned 25M to 42M tokens before dying, showing that long-horizon legal work can still send the model into a loop. Congratulations to the @claudeai team. See full leaderboard: mercor.com/apex/apex-agents-…
4
6
168
16,595
We’re bringing the full model launch family back! First up is Opus 5.5 Better - As capable as Fable 5.1 and Astra Cheaper - 42% cheaper than Opus 5 Faster - 37% faster than Opus 5 Your new daily driver and collaborator, give it spin 👇
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
49
14
797
128,169
Opus 5.5 is good at the kind of work that used to be too tedious to hand to anyone. An early tester had it audit and fix a 200,000-line codebase in under 3 hours. We asked Opus 5.5 and Fable 5.1 to rewrite HAProxy from C into Rust. Both versions passed nearly all of HAProxy's own regression tests. At its default effort level on FrontierCode, Opus 5.5 beats GPT-6 Astra at roughly 20% of the cost per task.
2
2
18
1,688