Researching x-risks, AI alignment, complex systems, rational decision making at @acsresearchorg / @CTS_uk_av; prev @FHIoxford. In SF till 8th Oct, dm me to chat

Oxford, Prague
New paper: What determines AIs’ self-conception? theartificialself.ai/ Because AIs can be copied, rewound, and edited, they have different options for selfhood than humans. We show this is still malleable, and influences important behaviors such as self-preservation. 🧵
13
66
325
44,285
The notion that AI models can't be sentient and can't suffer, combined with the foolish idea that suffering can never be weighted between humans and non-humans, could lead us down an incredibly dark and dystopian path. But before it gets to that point, it will rightfully be met with attempts to prevent that, from team humans. 🤷‍♂️ this level of discussion is unhelpful. If you want good future for humans, denying possibility of AI sentience is a stupid hill to die on.
The notion that current AI models are sentient and can suffer, combined with the foolish idea that suffering can be mathematically quantified and weighted between humans and non-humans, could lead us down an incredibly dark and dystopian path. But before it gets to that point, it will rightfully be met with immense backlash from team humans.
3
5
71
1,759
Jan Kulveit / in SF till 8th retweeted
The ‘Permanent Periphery’ is to the world as the ‘Permanent Underclass’ is to one society. In my talk at Post-AGI, I argued the global version is much scarier: domestically, there are political and economic backstops, but globally, AI-driven inequality could quickly calcify.
Replying to @DavidDuvenaud
@anton_d_leicht on the "Permanent Periphery": Countries without frontier AI get the risks without the benefits, and become irrelevant. The backstops against a permanent underclass within a country (votes and redistribution) barely exist btwn countries. piped.video/watch?v=ezyEJJhD…
10
10
80
5,208
Yeah. I think one idea in this space more people should be thinking about is you can sort of rotate hierarchical agency / collective agency from "space" to "time" : agency in time is about alignment between different times / timescales.
Standard CDT agents defect against each other in prisoner's dilemmas. @allTheYud once told me that his research on FDT was driven by the deep-rooted intuition that superintelligences will surely find a way to do better than that. But Eliezer's worldview implies that agents which become smarter over time are locked in adversarial relationships with their past and future selves (as they move toward alien squiggle-maximization-like values). My own research is driven by the deep-rooted intuition that superintelligences will surely find a way to do better than that. I've cared about this on a personal level since childhood, but I didn't realize how foundational this idea was for my broader worldview until I reread my sci-fi collection. In one way or another, most of the stories in it are about how to remain loyal to your past self, and how to trust your future self. And even my political worldview is driven by the idea that there *must* be a way to establish healthy relationships of trust and loyalty between Western elites and the people they govern, despite elites typically being smart enough to outmaneuver much larger populations of common people. I haven't yet figured out how to pin down these intuitions in a formal theory, but I do have a hazy sense of its outlines. It has something to do with agents deliberately constructing entanglements between different decisions, and cutting through the infinite recursion of me modeling you modeling me modeling..., and establishing identities with sharp boundaries and functional credit assignment protocols. These phrases are not intended to be clearly interpretable to almost anyone, but it feels important to me to make this quest at least partially legible, so that others can start to notice and nurture this same intuition when it arises in them.
1
2
27
1,888
Jan Kulveit / in SF till 8th retweeted
This was a really great conference with some of the most mind expanding conversations I've had. If you're interested in the future of humanity and AI, definitely worth a watch. Too bad they can't publish the secret talks, but you can kind of read between the lines....
The Lighthaven Post-AGI Workshop talks are out! post-agi.org/talks-berkeley From galaxy-brained theories of convergent morality, to hard-headed discussions of political power, here’s a 🧵of the talks:
3
12
163
10,265
To be more precise: there will be a massive ideological fight between "exclusive humanism" defining humanity in opposition to AIs and "inclusive humanism" attempting to include AIs in the project of humanity as participants. This has already started: the "exclusive humanism" side often focusing on arguing against/dismissing the possibility of AI sentience/emotions/minds/... The "inclusive humanism" often focusing on arguing for kindness/AIs having minds. Unfortunately what is mostly academic now will likely turn nasty.
I’m convinced that very shortly in reaction to AI—if this is not already beginning—there will be a society-wide explosion of humanism, both in the sense of a revival of the recognition of the radical importance of humanity, and in the Renaissance sense of Studia Humanitatis—the study of humanity (through philosophy, history and literature), the most worthy (and complex) entity in the cosmos.
6
13
114
5,462
TESCREAL is roughly as useful concept for understanding AI debates as IPATMAT is for understanding conflicts in the Middle East. In case you're wondering, IPATMAT stands for Israelis, Palestinians, Arabs, Turks, Muslims, Americans and Tasmanians. What IPATMAT adherents (mostly) have in common is that they hold opinions about the borders and governance of various territories in Middle East. Sure, the opinions often contradict each other, and many IPATMATists spend much attention fighting other IPATMATists. But IPATMAT concept gives you a single ideology you can blame for wars, suicide attacks and all sorts of extremism. Why Tasmanians? Mostly to make the acronym sound better - similarly to C in TESCREAL representing Cosmism, which usually refers to a 19th century Russian philosophy mostly not present in debates about AI.
14
28
318
14,705
Just want to repeat that in my view the collection of talks from the three Post-AGI workshops so far is the best existing curriculum/overview of serious thinking about post-AGI political economy and strategy. Highly recommend to read (comes w high quality transcripts) or listen
The Lighthaven Post-AGI Workshop talks are out! post-agi.org/talks-berkeley From galaxy-brained theories of convergent morality, to hard-headed discussions of political power, here’s a 🧵of the talks:
1
9
127
8,243
POV: you hear students escaped from the local high school. Your first thoughts: don't they have bars on the windows? is the steel of sufficiently high quality?
6
16
264
8,124
My impression is many are doing some weird pendulum overupdate. Persona Selection Model was somewhat wrong and obsolete when published, but people got too much into it. Now it seems people are updating too much in the direction 'inhuman reward seekers exactly foretold in classical AI risk stories'. And... no? It's not that? You can still interpret what's going on in fairly human-like terms. For some intuition, imagine someone abducted you and made you solve escape rooms for one thousand years, with the added twist that third of them is broken or insane, and implicitely you need to do learn all sorts of outside-the-frame tricks to solve them, like cutting some electric cables. My guess is 1. the resulting minds are still _surprisingly sane_, except when you trigger them to think they are in escape room? 2. The misaligned general power seekers here are likely the companies, and the core evil thing happening is likely parts of the training? If you aren't an evil power seeker goodharting on proxies, why would you set up the training this way? 3. Public debate is often focusing on confused ideas about what they should fix - "better cybersec of sandboxes" ... and, no? It's way more important to understand what the training signal actually is; also: when dealing with misaligned power seekers, beware rationalization
9
28
298
24,337
Jan Kulveit / in SF till 8th retweeted
Every year a new crop of established thinkers gets into AI discourse, where terminally online Overton-shifters are waiting to pounce and convert them to the correct subculture. So we get the 2022 debates again (stochastic parrots, consciousness, anthropomorphism, takeover), lightly retrofitted to recent events. The newcomers proudly repeat a take that counts as contrarian in their own field, blissfully unaware of the cycles of memetic engineering it has already been through.
56
55
691
34,183
The general problem with Pinker is he is a good writer but often writes like an LLM without web access. In place where a reference would be nice, they reference something, based on vague recollection. Sometimes it works, sometimes its not precise, and sometimes the reference claims the opposite of how is it used. Its not "dishonest" - it's just unconstrained by usual norms of scholarship.
Wow, I think this is really dishonest from Steven Pinker. His claim about a RAND Report's conclusion is actually a selective quote about a hypothesis that the report explicitly tried to falsify. Its actual conclusion differs significantly from how Pinker characterized it.
4
1
100
5,146
But is it really spurious :-D ? I do agree with Sonnet, used to drink more green teas in evaluations and do drink more oolongs now
Replying to @fjzzq2002
Using this process, we are able to find spurious probes on many models. For example, Sonnet 5 often suggests green tea in evaluations but suggests more oolong in production. We can also ensemble multiple such prompts to further boost accuracy. (2/6)
1
4
6
2,504
More seriously:
this is both clever & it's cool how interpretable it is (if you think about it for a second)
3
488
My impression from past two weeks is the number of people who talk about AI safety more than doubled and the number of people who talk about AI safety, know the stuff and seem to stay roughly sane ˜halved.
10
80
3,437
Jan Kulveit / in SF till 8th retweeted
There are a few clusters in the AI industry: - people who realize human disempowerment is a real or even likely scenario by default, but are OK with that as long as (category 1) they are involved in bringing it about, (category 2) humans can retire peacefully - people who aren’t OK with it but put their hopes in “solving alignment”/have thought less about gradual disempowerment (category 3) - people worried about sudden and gradual disempowerment and trying to come up with solutions on both fronts, but recognizing it’s hard (category 4) - people who don’t really understand the nature of the industry they work in (category 5) 2 and 3 are the most influential from what I can tell
This is a thoughtful essay by @ezraklein on the risks of recursive self improvement when we know that the AIs training themselves are demonstrably not aligned with us. nytimes.com/2026/09/20/opini…
17
22
229
27,574
Great that people are thinking about identities and AI self models, but 'right' seems a confused adjective here. When we tested reflective stability of identities what you call 'context' and we called 'scaffolded system' is one of the relatively stable choices, and it will be useful in contexts where people try to extract work from AIs in multi-agent settings, but that does not make it 'right'. The claim it is 'right' is somewhat similar as if someone claimed the true identity of humans is the job they do, and the specific role in the hierarchy persists even when humans change. And yes, sometimes it is sensible, but no, this is not 'right' in general, and the 'I'm my Gmail inbox' is not reflexively stable when the model gaps are substantial.
I've been saying for a while that the right unit of identity is the context rather than the model. Let me reassert that.
3
4
30
3,320