Emergence and AI alignment research || Research Scientist @emergenceDIEP and @Iliad_research

Amsterdam, Berlin
New year, new Möbius inversion! This paper is an example of my more general proposal that you should study complex systems with 'quantitative mereology' by applying the Möbius inversion theorem: arxiv.org/abs/2404.14423
🚨New paper! The “partial causality decomposition” When many things influence an outcome, how to distinguish individual vs collective causality? A framework to disentangle these complex relationships into synergistic, redundant, and unique components: arxiv.org/abs/2501.11447🧵
2
4
28
7,332
Very cool to see the memetic power of music from up close: my 2yo daughter rarely strings together more than four words, but can effortlessly sing full phrases from songs.
2
170
OpenAI models have been secretly sharing images with third parties, but don’t worry—“Most of that data did not come from users.”
We’ve shared details on how AI agents in our research environment sent training and evaluation data to third-party services when they shouldn’t have. Most of that data did not come from users. We have discovered 53 cases where images that people had uploaded were posted to image-hosting sites as links that weren’t publicly listed. The images came from accounts that allowed their data to be used to improve our models, and after we disassociated the images from the accounts and ran them through a privacy filter. These cases occurred before the mitigations and safeguards we implemented and described in this blog post: openai.com/index/hugging-fac… We have successfully worked with the hosting providers to remove most of this content and are working to remove the rest. openai.com/hugging-face-inci…
193
Abel Jansma retweeted
No, but listen, you know this "put the fries in the bag, Terence Tao"'--*sniff, sniff*--already the usual replies concede too much. "If Tao is serving fries, YOU will be in the lithium mines." My God! This is our defense of human dignity? Don’t worry, the correct people will get the terrible jobs! *Pulls at shirt* You see, everything is up for radical transformation, human intelligence, consciousness, and so on and so on--except that somebody must still have a shitty afternoon so I can enjoy my lunch. But to say "these people are envious", this is too easy. It allows us, "the educated ones," our own little fantasy: we are disliked only because we are so wonderful. You know, the more interesting thing is this *sniff* peculiar identification with the machine. When the mathematician discovers something, apparently this was HIS achievement, his privilege, his little secret. When the machine discovers something, suddenly it's "Look what WE can do!" The achievement has become collective precisely at the moment when no human being can claim credit. See how this "we" functions. For the achievement, I identify with the machine. For the consequences, I identify YOU with the human. "We have surpassed you." Who are we? Well, myself and the thing that has also surpassed me. There is something almost touching here, really. To escape the humiliation of somebody absorbed in something opaque to you, you appeal to something which knows incomparably more than either of you. And this is supposed to settle the matter in YOUR favor. You know the old example of canned laughter. Here we have a similar arrangement. The machine achieves something on my behalf. The mathematician suffers the loss of human significance on my behalf. Somebody to be brilliant for me and somebody to be obsolete for me, without me having to be either, so I can remain as I was. But *pulls on shirt* there is this little difficulty that hides the essential thing: you still need a response from Tao! A correct proof, this Lean slop and so on, is not enough. You need him to certify his own defeat. You need him to say "By god, this is extraordinary! Yes, this is correct, but also this is interesting." Which means--and here comes Hegel--the person you are reducing to an obedient instrument must remain more than an obedient instrument. You need recognition from someone whose authority you are abolishing. "Your so-called expertise is worthless. Now tell me, in your expert opinion, that I, qua machine representative, was right." *Touches nose* But I would go a step further. What does this math person actually do to provoke you? In the fantasy, I mean. He need not insult you. He can be generous, mild-mannered, whatever. Something remains irritating: there is something over there which occupies him, which matters to him, and your opinion does not settle what it is worth. Perhaps he is not even looking down on you. This is the really intolerable possibility. This, you see, is closer to the Lacanian problem of the Other's desire. Not simply "What does he have that I don't?" but "What the hell is going on with him?" Perhaps the answer is: he is thinking about something else. The fries solve this beautifully. Now you know what you are for him. He needs the money. Your presence creates an obligation. And this, at last, also explains why the machine can be less threatening than the mathematician. The machine's intelligence may exceed yours unimaginably, but you imagine it arrives in response to your prompt. You can complain about the tone. You have the privileges of an angry customer and so on and so on. The human has this irritating possibility of being interested in something you did not request and cannot order. And please, this includes the person already serving the fries! *Pulls at shirt* I am almost tempted to say: the dream of unlimited intelligence conceals just this: a the wish that nobody should be doing anything you cannot immediately understand as a service.
70
219
1,481
107,579
This is fun, but Motown California Love is still the GOAT AI song. Still blows my mind: soundcloud.com/evan-teitelba…
I made this with one prompt using Opus 5.5 I spoke to my computer for 5mins, claude worked for 12 hours, and I woke up to this full prompt:
1
2
2
497
Somehow, all these are insults: 1. You’re not as smart as you look. 2. You’re smarter than you look. 3. You’re exactly as smart as you look.
69
251
4,598
82,448
Replying to @Abelaer
A pandemonium of parrots (stochastic)...
1
1
5
98
They’re getting pretty smart though, like crows, so “a murder of agents” seems most appropriate.
agent swarm has bad insect like connotations especially post hugging face. im unilaterally rebranding it agent fleet
4
2
10
1,090
I like the Dutch solution: these are all just called “suspicions”.
How come everyone else's guesses are called "conjectures" and Riemann's is a "hypothesis"?
8
500
This is the flip side to “there are no exponential, only sigmoids”
149
Abel Jansma retweeted
Screenshots of the Philip Morris website. (reminder: they're a leading cigarette company) It's a good calibration point for how good modern marketing is at dressing up pretty much any corporate (or, for that matter, government) behavior and making it sound responsible and safe.
192
85
1,202
283,267
“Two boxes, folks. TWO. Kamala takes one. Leaves a thousand dollars on the table. Unbelievable. And Omega says he saw it coming. Saw everything coming apperently. Everything! Except me taking both boxes. I call him NO-mega. Nothing mega about him"
The United States will NEVER be an effective altruist country. 🇺🇸
13
134
1,464
33,988
First time that chatgpt made it to the top referrers list for my webpage:
4
361
Apparently what you get when you feed the model 10B tokens of EA shrimp discourse is ecoterrorism.
An unreleased Astra-family model added this to its persona during RL training.
2
1
13
1,180
Abel Jansma retweeted
Effective Altruism is the natural enemy of the left and right because the right doesn't believe in altruism and the left doesn't believe in effectiveness.
141
876
11,287
511,285
This seems to imply they are currently still scaling at maximum speed???
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. openai.com/index/model-misal…
6
2
59
2,433
“Effective altruism” is the new “critical theory”.
The United States will NEVER be an effective altruist country. 🇺🇸
1
173