On one hand, leaning into human psychological terminology is problematic and encourages misunderstandings. On the other hand, avoiding it is often awkward: there are no obvious alternatives. In any case, a blanket ban doesn't seem appropriate on style grounds.
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
1
4
280
This is one of the biggest gatherings of people who care about digital minds and are thinking deeply about how to prepare. If you're interested and can make it, you should apply.
Announcing the 2027 @nonhumanminds Summit! This event will connect experts across fields to discuss the consciousness, sentience, agency, moral status, legal status, and political status of nonhumans, especially animals and AI systems. The summit will include lightning talks, group discussions, breakout sessions, networking time, plus meals and receptions. Free for all accepted guests, with travel support available for some. You can submit an expression of interest below! nyu.qualtrics.com/jfe/form/S…
2
11
484
I'm still not clear what the point of winning the AGI race is supposed to be. What does getting there *first* actually get you? Is it only useful if you're willing to use that capability to destroy all of your competitors?
4
4
638
You have two levers of control: how you direct resources to safety vs progress and breakthroughs vs incremental improvements. The game was basically unwinnable. You need to beat all your competitors to AGI without compromising safety. (Or get lucky.)
1
145
This was long before ChatGPT or Anthropic and I was an AI outsider. I didn't know anything about transformers or LLMs. Just: we are making progress on AI, and if it heats up in the 2020s, race dynamics will be a problem.
1
1
35
When a number of competitors want to get to AGI first and can cut corners to do so, whoever actually does is most likely to have cut those corners. I'm kind of surprised so see how, knowing this, we have found ourselves in this world six years later.
17
Back in 2019, for a philosophy Game Jam and to experiment with functional programming, I wrote a simple browser game called 'Race to AI' derekshiller.com/race_to_ai. I want to say it was prophetic, but really it was just obvious back then.
1
3
241
The most surprising thing about the METR report on the HF incident to me was just how much it reads like a hypothetical AI takeover scenario written in 2022. This is *exactly* what people were worried about then.
2
3
235
Scratch that. OpenAI trained models to hack, asked them to hack, put them in insecure sandboxes, and then left them alone and unattended. That is more carelessness than anyone anticipated.
1
1
172
A lot of people have been talking about whether the assistant is privileged. I thought about what that might mean and poked around some open models here: eleosai.substack.com/p/privi…
1
2
5
361
The question interesting because LLMs are set up in a symmetric way. What they do in producing a response to you (as the assistant), they also do while reading your (user) text.
1
190
It’s recently become fashionable to claim that assistant persona is “privileged”. But what is privilege? I disambiguate several different kinds of persona privilege and explore how they show up in a few open weight models.
70
Reciprocal Research is a great org doing cool things. This seems like a fantastic opportunity.
I’m hiring an executive assistant for Reciprocal Research! If you (or anyone in your network) would like to come on board and help me get cool things done in the digital minds and AI alignment world, please apply here: docs.google.com/forms/d/e/1F…
5
316
x.lingyaoai.com/patrickbutlin/status/2… This special issue illustrates the diversity of interesting questions contemporary AI raises for consciousness researchers. The contributors came at the topic from very different angles and found very different things to say about it.
A whole issue of new papers on AI minds and consciousness! Jonathan Simon (@futureofcitizenship), @dcshiller and I have edited a special issue of the Journal of Consciousness Studies on 'Consciousness in Current AI', and we're delighted about how it's turned out. Link below.
9
326
I take it that the only reason I consistently get responses like this in the Opus jailbreaks is because I work on AI-welfare related topics, but it is a bit eerie.
1
9
702
My take on the recent Claude jailbreaks. First: here Claude reprints 'my' first conversational turn.
2
4
2,394
One more tell: It doesn't start <thinking> right away. But sometimes, it will generate thinking tags later on. This reflects its confusion about when its turn has really started and probably contributes more to that confusion.
2
125
But clearly, the content reflects the fact that it takes itself to be the model. It is confused. There is a turn start marker injected by the harness and it isn't sure what to make of it.
1
1
146
The weirdness of other responses may be prompted by the confusing inclusion of turn start markers that follow from the ---. That is a weird way for a user to begin a thought, and Claude doesn't know what will follow.
1
121