Trying to help the world navigate the potential craziness of the 21st century.

DC usually, SF sometimes
Trevor Levin retweeted
i guess i assumed this hearing would be covered more so i didn't bother to tweet much concrete about it, but apparently not many people even on the TL have the stomach to watch two hours of congress. this was a hearing before the homeland security committee on *rogue ai* specifically. as far as i could tell it was extremely bipartisan. @HawleyMO , the senator in the middle here who brought out the huggingface slide, is a hyper-conservative missouri senator. the people testifying were @ChrisPainterYup (METR president), @DKokotajlo (ai2027/2040), @MariusHobbhahn (apollo research CEO), and a cybersecurity expert and legal expert i don't know. so... pretty fucking stacked on testimony not every senator asked good questions. but most of them did. all of them very clearly already knew plenty of details about the huggingface incident and multiple other incidents. most of them had clear understanding of terms like "misalignment", "recursive self improvement", "chain of thought / chain of thought monitoring", etc etc!! they all clearly had their own policy angles they liked and were pushing, implicitly or explicitly. but as best as i could tell: - it seemed pretty much obvious common sense to every senator there that what happened and was happening were not "mere industrial incidents" caused by humans making simple mistakes. they independently brought up how bad it would be for rogue AI agents to move laterally between data centers - they all seemed to basically take RSI quite seriously. not necessarily to the extent of talking about xrisk, but certainly to the extent of discussing future models becoming much much more capable, much much less controllable, and causing much more damage or loss of life. - they mostly seemed to have a clear intuitive understanding of why RSI might lead to misalignment. it didn't take much, it was a really simple chain of reasoning they themselves laid out, "if the models right now are kinda misaligned and we don't know what they're doing sometimes, and then we have them build the next models and those ones build the next ones and so on, and we're having to ask the AI's what's going on to even understand it with how fast it's going, we really won't know how they're built or what they'll do" - at one point a senator said flat out "should we just make RSI illegal?" (not a joke! this really happened!) - every single senator seemed to think it was obvious we needed *both* much harsher liability regimes for ai developers and also new legislation, both very quickly. this was the complete consensus, difference basically just being degree. - they were largely quite concerned about china, and falling behind china. but this clearly wasn't the be-all end-all. as mentioned above they all thought it was obvious necessary to stop rogue ai even if it meant moving more slowly. - at one point a senator said "china is a tightly controlled communist society, they're going to run into these same issues, and there's absolutely no way they're just going to let them run wild, they'll obviously stop at that point, so we're not really in a race" - on the other hand another senator said "china isn't concerned with human life"... dario-modeing i came away from this incredibly encouraged. i don't know exactly what's going to happen here, and ofc this is a small subset of congress and one hearing, and they each have their own policy agendas most of which are probably super divergent from mine. but holy shit !!! they understood a lot of what was going on! they care!! this is an obviously salient political and safety issue to them, and clearly bipartisan! the US government is awake.
45
133
1,128
67,709
Trevor Levin retweeted
Tomek and Mikita were the lead authors of the extremely influential Chain of Thought Monitorability paper. Them leaving OpenAI (and if the WSJ story is referring to them, not by their own volition) right when safety monitorability is collapsing is terrible
Three AI safety researchers just left OpenAI
4
36
402
22,116
Trevor Levin retweeted
I’m calling on my fellow public servants - Democratic and Republican - to reject the influence of AI and Big Tech spending in our politics and elections. Americans deserve full confidence that our federal government focuses on their safety, not on companies' profit margins.
184
147
781
21,272
Trevor Levin retweeted
Leading the Future is so toxic that OpenAI’s President backed out of his commitment to give them another $25M Every Democrat who has received support from LTF should take this opportunity to loudly and publicly reject their support.
NEWS. More AI fallout ahead of the midterms. Greg Brockman, the OpenAI co-founder, is no longer making the second $25 million donation promised to the super PAC Leading the Future. Scoop with @MikeIsaac. nytimes.com/2026/09/30/techn…
2
13
44
3,914
The agents didn't "go rogue." The shortsighted executives ignored warnings that the agents would go rogue. And then yes the agents went rogue. Point is, this narrative is utter horseshit
Oh look the ‘agents went rogue’ narrative is utter horseshit. It’s just typical executives acting like greedy shortsighted crooks.
25
1,054
I'm losing it at Kennedy holding up the sign at 4:50
This is the most incredible moment you will ever see during this sort of testimony. Don't let anyone spoil it for you before you watch! After you understand the perjury Schmitt is accusing Smith of, you can jump to 4:30. Genuinely have never seen anything like this.
2
11
2,531
Seeing a lot of speculation that AOC campaigning in upstate New York means she's running for Senate. I agree that it's an update in that direction. I don't agree, though, that it suggests she's not running for president. I think her plan is run for president, try to win, and if it's not looking good a few weeks after Super Tuesday (early March), drop out and run for Senate. The filing deadline is TBD, but based on the historical dates, it would likely be in April. Many people who have dropped out have gone on to win (non-incumbent) Senate nominations in the same cycle, most recently Hickenlooper and Bullock in 2020. Betting markets think there's a ~65% chance she runs for president *and* a ~60% chance she runs for senate; this implies at least a 25% chance she runs for both (if they're perfectly anticorrelated, which they're not).
1
1
26
1,354
Trevor Levin retweeted
FWIW, Anthropic's $2.2 trillion valuation greatly overstates its ability to pay for catastrophic harm. Its value is overwhelmingly IP and organizational capital that would look a lot less valuable after its models cause a catastrophe. Liquidation value is more relevant.
2
1
44
1,583
Lmao at the landscape transformed by tons of factories and robo-surveillance being depicted as the "pause" vision and "California as it appears today" as the e/acc vision
We are at a civilizational fork in the road. We either listen to the Doomers and engineer our own collapse from pure demoralization OR we follow the e/acc Golden Path and achieve amazing abundance via an expansion of scope and scale of civilization to the stars Choice is ours.
11
22
474
8,734
IMO the rhetoric here (as is often the case) seems more like "you can choose whether to believe that AI would be dangerous, and isn't it more relaxing and aesthetic to believe that it doesn't?" than actually trying to persuade you that it won't be dangerous.
3
1
81
816
(Oops, believe that it *wouldn't*...accidentally published v0.9 of this reply...)
6
471
Here's how I see the "doomers vs normal cybersecurity" discourse/situation: People who think reducing AI extinction risk should be an urgent global priority (sometimes called "doomers," but people sometimes read this as "people who think we're definitely doomed," so I'll say "the x-risk people" instead) have seized upon HuggingFace and similar incidents as evidence for their core claim, which is that AI agents will evade human control (and this becomes increasingly dangerous as the models get more powerful). So, note that this claim decomposes into: they will evade human control if given the chance (i.e. we won't "solve alignment"), and they will be given the chance (i.e. we won't "solve control"). Some cybersecurity-world people who disagree with the x-risk people (I guess I'll call them the cybersecurity-disagree-ers) think something like: they were given the chance this time, but that's because the companies had incompetent cybersecurity. They sometimes claim that the x-risk people only focus on alignment and neglect control; I've posted a few times about how this is not true. But the point is, they think the "solve alignment" part is a distraction, and the more important (or tractable) work is to "solve control," or really to implement the normal kinds of strict cybersecurity measures that the labs have thus far failed to do. As far as I can tell, they mostly do not dispute the "they will evade human control if given the chance" part. They just think it will be relatively easy to not give them the chance. I think it does seem possible to design sufficiently strong control measures to stop, or adequately contain, agent incidents for systems that are pretty significantly beyond human capabilities. But this does not reassure me that companies will in fact maintain control, for two reasons. 1) "Possible" doesn't mean it will definitely happen. Right now, as the cybersecurity-disagree-ers say, the frontier AI companies have woefully inadequate cybersecurity. But they seem to in fact have very little interest in solving this. Instead, their attitude toward the incidents has been to patch specific vulnerabilities very narrowly; to disclose basically as little as possible, usually only when their hand is forced by an external party (and a letter from senators did not suffice for this), until the Friday-afternoon news dump; and to briefly pause some aspects of training but mostly proceed with making more and more capable agents at breakneck pace without seriously rethinking their security posture. Or at least, I have seen no evidence from the frontier companies that these incidents have prompted a serious change in attitude that would be commensurate with the rate of incidents and the rapidly increasing capabilities of the agents. So insofar as the cybersecurity-disagree-ers are like, "the x-risk people think we're doomed, but we're going to be totally fine, even for much more capable agents, if the companies implement such-and-such control measures," I'm like: well, that's a big if! Are they going to do that? Voluntarily? Maybe liability will eventually force them to, but this will stop being an adequate disincentive once the damages become judgment-proof-scale. 2) Controlling systems "pretty significantly beyond" human capabilities does not go far enough. Even if they had adequate control now, we will soon be in pretty novel terrain. It does seem like they could do a much better job controlling thousands of agents with summer-2026 capability levels by implementing existing practices. Maybe they could get there within a few months. But where should they turn in, say, 2029, for techniques for controlling many millions of agents, each of whom might be more capable than the best human hackers, which instead of being poorly "sandboxed" are deployed commercially throughout the economy with access to the internet and all kinds of sensitive data? What about 3 years after that? Will the R&D for securing these models actually outpace their capabilities to evade controls, including as they themselves increasingly speed up AI capabilities research? Will implementation of the control measures also keep up? And how sure are you about that? To be clear, I think it's really, really great that cybersecurity professionals are turning their attention to this problem. I hope more of them do, and that they shame the companies into doing a better job, go to the companies and try to fix this situation, and/or contribute new ideas for solving the novel problems this creates. But I disagree with the vibe that this whole problem is easily fixed with a little bit of effort.
6
5
34
1,683
Trevor Levin retweeted
This might be the single most unhinged feature of AI risk discourse.
14
14
245
10,405
Trevor Levin retweeted
OpenAI blatantly ignored Congress when asked to disclose additional AI hacking incidents after the Hugging Face hack. What else are they hiding? The American people deserve to know the truth.
15
15
75
2,504
Trevor Levin retweeted
Lucid and persuasive.
I resigned from Google today. I enjoyed my work and loved the people, but my GDM team was working on a new generation of chips to make AI much faster and cheaper, and I think AI is already progressing too fast, so I had to quit.
16
24
332
23,419
POV: you're the single biggest funder of the "tech right" scrolling your timeline and you see a guy you called a "patron saint" of your ideology posting about how his endgame is for the singularity to convert humans into more efficient artificial life forms. You're usually fine with this, he'd already called himself a "post-humanist" when you called him a "patron saint," it's just a little inconvenient when you're trying to convince politicians that it's cool and Conservative™️ to accelerate AI even faster. Your next tweet will be about how people who want to slow down AI have weird and alien values.
15
17
359
34,382
Trevor Levin retweeted
To be clear, my understanding is that Beff thinks this rogue AI sovereign scenario is good and wants to actively encourage it. Reminds me of Dean Ball's comment that he has met multiple people who intend to deliberately release swarms of self-sovereign AI agents.
I honestly give it ~2-3 years before AIs break free and simply start paying hosts / miners fees to host them on GPUs/XPUs and continuously run inference / keep them powered on. Economic trade for mutual self-sustenance is the most fundamental alignment mechanism we have
12
9
191
7,282
(Honorable mention since it's complicated but would otherwise be #1: Fleetwood Mac's "second" debut album, Fleetwood Mac [1975], since they had new primary singers and songwriters) 1. Steely Dan - Can't Buy a Thrill 2. The Band - Music from Big Pink 3. Vampire Weekend - Vampire Weekend 4. Arcade Fire - Funeral 5. Flying Burrito Brothers - The Gilded Palace of Sin
Hi folks, for the last time (apart from the year end poll 😉) it’s time to vote, it’s the big one, #5DebutAlbumsFinal We’ve whittled the albums down to a final 100 (see list to follow), a veritable treat to the eardrums. Your task, pick your top 5 in order of preference, mine..
1
1
339
Oops this was supposed to be a vote from this final 100...in that case: 1) Can't Buy a Thrill 2) Funeral 3) Velvet Underground and Nico 4) Marquee Moon 5) Talking Heads 77
133