Lance Martin retweeted
Fantastic conversation - Joe asks tough questions! Did we answer them well? I think this might be the first appearance of @the_marwell in public - he is a complete legend within Anthropic, an unbelievable amount of critical programs simply would not have progressed without him. Getting to know him via this podcast is crucial lore.
Some of my friends will be mad I recorded this, and comms people at Anthropic objected to releasing certain parts (and delayed it). But it's important for leaders to make these conversations happen. And I have a lot of respect for these two. Over a month ago, I sat down with Anthropic's key technical leaders @_sholtodouglas & @the_marwell for an optimistic insiders’ view of the AI frontier, with hard questions too: open source, regulatory capture, slowing down USA vs China, and more.
21
28
448
50,802
Lance Martin retweeted
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
1,424
5,859
50,780
6,403,152
Lance Martin retweeted
Today we launched Claude Sonnet 5.5! 30% faster and costs up to 30% less than Sonnet 5 for most work. Strong for well-scoped everyday tasks (fixing bugs, creating docs, slides). @RLanceMartin had models paint the same photo in code. You can see the jump:
83
49
1,364
101,501
i tested sonnet 5.5 on "code-to-painting". every pixel is generated by python the model wrote. it's shown a photo, writes a brush engine, renders, looks, and revises. you can see the large step-up from sonnet 5 in terms of coding + visual reasoning.
9
10
121
12,780
credit to @jkeatn for ideas related to code-to-painting and @IceSolst for the reference image!
for the past few months i've been asking our models to paint. opus 5.5 is very skilled at emulating different styles every image here is a python program generated pixel by pixel. there is no image model, and no off-the-shelf art software. instead, it's about 7,500 lines of code using standard libraries to emulate different brush styles. the agents don't use any pictures as reference, instead working only from what they know about each painter
1
15
14,104
Lance Martin retweeted
Small change, but "/claude-api prompt-audit" is now also "/checkup prompt-audit" The name made it sound API-only, even though it always worked on your Claude Code setup. It checks your CLAUDE.md, skills and agents for instructions your model doesn't need anymore. Really useful!
91
111
1,666
169,292
Lance Martin retweeted
Replying to @buildwithrajath
yeah, we've had a few fumbles the past few months we're going to listen to you more, incorporate your feedback more, and we want to talk to you more (DMs are always open) on that note, go fucking build.
47
12
715
100,684
Lance Martin retweeted
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more: anthropic.com/news/claude-di…
1,562
5,315
40,998
25,449,795
useful tip for Opus 5.5: run “/claude-api prompt-audit” in Claude Code. this checks you skills, agent.md, Claude.md, prompts and removes anti-patterns that hobble frontier models. i updated the skill w/ the latest Opus 5.5 guidance.
102
199
2,888
287,660
Lance Martin retweeted
opus 5.5 = personality of opus 4.6 that we all *desperately* wanted back + the intelligence & taste of fable 5.1. and a 25% usage bump! i honestly don’t really see a reason to use another model right now? s-tier release.
138
147
5,721
197,551
Lance Martin retweeted
for the past few months i've been asking our models to paint. opus 5.5 is very skilled at emulating different styles every image here is a python program generated pixel by pixel. there is no image model, and no off-the-shelf art software. instead, it's about 7,500 lines of code using standard libraries to emulate different brush styles. the agents don't use any pictures as reference, instead working only from what they know about each painter
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
193
343
5,024
825,193
Opus 5.5 is out! excellent at coding, great at writing, reduced pricing. run '/claude-api migrate' in the latest Claude Code to update your application code. run '/claude-api prompt-audit' to ensure your skills + prompts are well-tuned. short video overview:
15
25
380
35,536
Lance Martin retweeted
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
3,342
9,034
97,087
27,987,486
Lance Martin retweeted
In light of the progress in mathematics, we at Edison Scientific and FutureHouse have assembled a set of Millennium Problems for Biology. They are chosen to be very hard to solve but very easy to validate in a simple laboratory environment. Any of these, if solved, would mark a major advance in biotechnology, and most of them would contribute materially towards curing disease. These are, in some sense, the “last reasonable eval” for AI in biology. This was work primarily by @MichaelaThinks and myself, with contributions from many others. Short descriptions below. The full descriptions of the problems with acceptance criteria are at the Bio Millennium Problems website, linked in the next post. Share more if you have ideas. If they meet our criteria, we’ll add them to our list (with attribution and permission).
172
586
3,486
508,585
Lance Martin retweeted
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
2,060
2,643
31,350
5,594,617
Lance Martin retweeted
Dario's essay points towards the right path forward. The details need working through, but the direction is correct for meeting this critical moment. This is also why we recently put out our proposal for an industry-wide standards body for frontier AI.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
820
945
9,071
1,411,481
Lance Martin retweeted
I personally believe there is a much greater than 10% chance that AI relieves the burden of disease and ushers in a post-scarcity Star Trek future of abundance and unimaginable human flourishing and scientific discovery. I'm very grateful to have colleagues who have been thinking deeply about alignment, and working on it for a very long time, to prevent bad futures and make sure we get to good outcomes for humanity.
62
73
699
48,384
Lance Martin retweeted
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
5,174
7,195
67,617
17,078,035