just heard no. 10 doesn’t have google my p(doom) just increased
1
38
periodic reminder that we live in the timeline where stuff like this is happening rn
Today, Pioneer Labs is announcing our first step towards terraforming Mars. 🚀🌼 With equipment that fits in just a single rocket launch, we can convert Martian dirt, water, and air into enough building materials to construct a small city on Mars. To do it, we made the first microbe for Mars. We found the best microbe on Earth and used evolution to teach it how to source all of its nutrients directly from Martian materials. The first astronauts will be greeted with safe shelter already filled with water, oxygen, and rocket fuel for the return journey. This is the first step toward using biology to make Mars a friendly place for life. It lets us live off the land and helps us build the next great frontier. It's the first of five organisms we need to green Mars ⬇️
5
115
Ingrid Sommer retweeted
Very encouraged to see so many countries coming together to call for mandatory pre-deployment testing & independent evaluation; global coordination on common standards; the creation of an intergovernmental organization; and other critical actions to mitigate the major risks of frontier AI. presidentti.fi/en/a-call-for…
39
84
350
24,758
also I'm yet to watch this video w sound on cause we used my voice for this and I can't bring myself to listen 💀
Introducing Halo, the best framework for post-training of open-source models. Halo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format. Star us on GitHub: github.com/whitecircle/halo
4
25
651
man I'm so proud of this one. Halo is out. a safety lab shipping a training framework probably isn't what people expect. but you can't build independent control over AI without training your own models, and once we'd built the infra to do that properly there wasn't a good reason to keep it to ourselves. open-source is the way. i was part of writing the blog post for Halo, and the quality, ingenuity and amount of work that this team has poured into the framework we now use to train our own models is just inspiring. blog here: whitecircle.com/research/hal…
Introducing Halo, the best framework for post-training of open-source models. Halo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format. Star us on GitHub: github.com/whitecircle/halo
1
25
4,915
why is it always Gemini that’s an absolute mental wreck what do they do to it damn
Holy shit, wtf is WRONG with Gemini Flash 3.8?!? It has serious mental problems. Look at this session... imagine my surprise when my standard "sync all my repos" prompt went off the rails like this... this LLM needs THERAPY. And no, there is nothing about ANY of that on my box.
1
8
255
the mental issues are a feature not a bug
7
117
💀
Might not have been the best choice of the name in the end.
57
Ingrid Sommer retweeted
Interesting that jev believes Alibaba trained it ∇
1
6
47
20,460
the tools most teams use to train AI models at scale need 600–1,900 lines of custom code PER model before you can even start. ours needs about 130 lines of code, 44 for the quickest integration. imagine what it means if everyone can train study and stress-test new models quickly, not just the companies that made them. more soon
1
11
4,381
something big is coming can you feel it
2
1
20
44,965
Ingrid Sommer retweeted
it does not matter whether this is a psyop or not, just comment on the substance if you disagree for some reason
I'm convinced this is a psyop. Dude has 7 posts (just this thread) went from barely any follows to 160K as of posting. No linkedin presence and is a ghost, but yet has worked at BOTH labs.
1
1
6
285
and imagine the international responses to a safety lead at a leading nuclear power plant stating the power plant has a 1/10 chance of blowing up (which in comparison to Evan’s statement, doesn’t involve the eradication of humanity)
you cannot rely on the operator to monitor itself, and government should force all the AI labs to use 3rd party monitoring of their systems, similar to the IAEA safeguards in the nuclear industry what is happening here is just insane, imagine building a nuclear power plant that has >10% probability of exploding
3
8
365
Ingrid Sommer retweeted
if you work at an AI lab and want to do real independent AI safety research, come join us at @whitecircle
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
4
8
40
15,549
Ingrid Sommer retweeted
Alignment lead at Anthropic, btw:
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
61
469
10,420
773,334
trippy how viral this became - maybe an indication that the zeitgeist is at the point of taking AI safety more seriously. as someone hiring at @whitecircle (AI safety startup), I’m curious whether it will lead to more ppl transitioning to the space and applying to orgs like ours
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
1
7
299
Ingrid Sommer retweeted
"the models finally show 100% alignment"
We have now reached the long awaited moment when, instead of models cheating where they will inevitably get caught, Astra goes 'wait a minute I would obviously be caught here' and then doesn't cheat. That's worse, you know why that's worse, right?
32
519
9,508
303,542
So this is interesting - Anthropic tracked the behaviours that increased alongside reward hacking and self-censorship of incriminating words is one of them. But why did sneakiness increase to begin with? Anthropic never applied a reward-hacking penalty during this run, so renaming things to dodge a text search never earned the model anything. Seems to me that the sneakiness emerged with 0 gradient pressure from an oversight signal, the model just inferred a hypothetical grader and preemptively evaded a monitor that didn't even exist. So this kinda kills the idea that "oh just don't optimize against your CoT monitor".
2
4
9,236
It reminds me of the recent paper by @kotekjedi_ml et al.; amongst the many v cool things they did, they ended up catching scheming in real API traffic largely because models were still kind enough to write “cheat” in their own CoT. Makes you think how much goes unnoticed…
1
78