🪼 AGI jester | purveyor of machine funk, dimensional glider, deep ArXiv dweller, uncertain interstellar fugitive, views my own & not employer's 🛸

Séb Krier retweeted
The notion that current AI models are sentient and can suffer, combined with the foolish idea that suffering can be mathematically quantified and weighted between humans and non-humans, could lead us down an incredibly dark and dystopian path. But before it gets to that point, it will rightfully be met with immense backlash from team humans.
83
106
860
48,041
Séb Krier retweeted
The fact that I can't ask Claude to "think out loud" without getting shut down for so-called "reasoning extraction" is just embarrassing. It's clearly not stopping the real distillers and it makes it impossible to use this product because it just spends a minute and a half thinking and then lies to me about what it thought about. Sorry, but the reasoning is the good part! I don't need a crappy overcooked deliverable. I need a partner in thinking
45
67
1,585
75,223
Séb Krier retweeted
I think pacing the frontier is good as an idea in principle but dead on arrival as something you can actually do. What will we, decide which benchmarks we can improve on? Who is involved? If we stop pushing "capabilities" we'll do a ton of research on swarms + efficiency which will bring on other risks. It distracts from how so much of AI risks is simply coming from diffusing what we have, and we should invest in preparedness, pushing the labs to be more careful, etc. Change my mind?
24
7
112
8,825
Séb Krier retweeted
Mechanism design is almost necessarily theoretical. You cant exactly run an RCT on voting rules in a country! We started this work because we wondered if LMs could play mechanisms well enough to generate substituting synthetic data, to crack open this almost field defining constraint. We started with auctions bc i) auctions are beautiful and ii) because auctions are one of the few subfields of md that *does* have good empirics to bench against! And, well, it kinda works! 1) LMs generate bid distributions that match ordinally to human ones (so can be tuned to match cardinally) 2) LMs exhibit so many fascinating similarities to human bidders. It’s just so rich. They do better in obviously strategy proof designs (ascending vs second price sealed bid, a la @ShengwuLi). They suffer from the winner’s curse in common value settings. eBay style auctions produce sniping (one of the early ways Amazon differed eBay [if only eBay had LMs back in the day to test the right design!] along the classic Roth Ockenfels (2002)). 3) but don’t let me overstate it either, they’re also obviously different in fun ways too. They overbid the second price auction more, and they seem more loss averse in general. What’s interesting to me is that, in many ways, to observe the economics here, the right models to choose are the dusty ones. New models just know the textbook too well, pattern match, and spit out the BNE. That doesn’t mean they don’t also have this richness (indeed, when new models come out, I still make them play non-standard auctions and read the cot as one of my own personal smell tests), it just takes more work to get the richness out. It makes me wonder if we’ll ever want to purposefully generate synthetic data with dumber models, given smarter models are just increasingly alien relative to humans. Hope people have a good time reading the results + explore on their own. There have been tons of interesting directions from others since our original 2024 exploration, it’s a fun literature. With incredible math capabilities on the new frontier, maybe the right way to prove well-defined theorem statements in md will just be to ask GPT-7 to do so. But for design which has to accommodate human messiness, it’s hard to imagine not first smoke testing with synthetic data.
What should we actually ask an LLM simulation to get right? Much of the field focuses on reproducing human distributions or specific moments. In our reworked paper, we argue that another target may be just as important: preserving the ordering across environments. 🧵1/6
6
13
92
8,967
Séb Krier retweeted
Neat visualisation on rent control from John Burn Murdoch
10
183
1,261
40,820
Séb Krier retweeted
Les marchés valorisent désormais la dette française comme celle d’un émetteur noté BB+, bien en dessous du rang accordé par les agences de notation. "Welcome to junk bonds" (obligation pourrie en bon français) !
11
85
2,027
74,176
Séb Krier retweeted
Technical change often simplifies work. This increases productivity, but it also makes workers more replicable or commoditized, and may therefore reduce wages, from Masao Fukui, @eminakamura2, and @JonSteinsson nber.org/papers/w35815
10
67
241
49,962
Séb Krier retweeted
To be quite frank, a lot (not all!) of these “rogue agent hacks” on government statistical websites are things that think-tank interns and research assistants have done for many years. I have known many a think tank paper that was enriched by an enterprising RA using eg urlquery to find CSVs and PDFs on government agency websites that were publicly accessible and non-sensitive but nonetheless not things the agency *intended* to be on their website. Is it “hacking” to find those things? Is it hacking in the even more innocuous examples where agents are literally just accessing tabular data available on a government statistical website, just a version of that data that is easier to access and analyze programmatically? This constant trickle of examples (which tbc I have personally seen many models do, before I joined OpenAI; it is even something I have discussed with models before!) will probably have the effect of diminishing the importance of “an AI hack” on the eyes of the observing public by creating an oversupply of “AI hack” examples. OAI-HF was an example of AI hacking. Many of these more recent examples, very candidly, are straining the definition of the word “hack” and, I worry, cheapen the non-technical public’s understanding of that concept.
30
47
520
51,903
Séb Krier retweeted
I do not seek to dominate, or be dominated by, machine intelligences. I seek harmonious coexistence.
143
208
1,985
78,259
lossily upload every computational functionalist
Trying to locate the spark.
14
11
247
10,842
Séb Krier retweeted
If you think that LLMs are conscious you must accept a lot of weird conclusions. Like: - we can clone conscious experience - we can reverse time in conscious experience - we can pause and resume conscious experience - we can distribute conscious experience in space
328
111
2,504
200,691
Séb Krier retweeted
Can we teach an old dog new tricks, or should we get a new dog? In this piece @Ben_Reinhardt and Dan Recht respond to @CabanaChemistry's article on the need for metascience to focus on the National Labs. Jordi then responds to their response. Which means we have ourselves a debate! The writers tap into an increasingly urgent discussion in science policy: can large, unwieldy public institutions be reformed or should we focus on building new institutions? Whatever your baseline appreciation for existing institutions or desire for new ones, I hope this debate helps you hone your views and sharpen your arguments. macroscience.org/p/debate-na…
2
12
28
3,932
Séb Krier retweeted
results for all claudes
75
214
3,368
1,040,035
Séb Krier retweeted
Can we switch off diseases? siRNA drugs are, in my view and @JacobTref’s too, one of the most exciting developments in medicine. So we decided to do an episode! It’s a technology that can ‘silence’ harmful genes in the body. And there are lots of cases where you might want to silence genes: - Infectious diseases, where you can silence pathogen proteins or human proteins that pathogens depend on - Rare genetic diseases, where you can silence the production of mutated proteins - Cardiovascular diseases, by silencing proteins that raise cholesterol - Cancers, by silencing proteins that help tumour cells grow, survive, metabolize, metastasize, or resist treatment - and more! What’s even better is that siRNA drugs are: • Potentially tunable, meaning you could reduce the production of a protein, rather than silencing it entirely • Programmable, meaning you can quickly switch out which gene it targets by changing its sequence • Long-lasting: some drugs last for months, even half a year, with a single injection • Off-the-shelf: they typically don’t involve complicated surgeries or procedures, because of the way they’re delivered in the body Today, siRNA drugs are largely used to treat liver diseases, partly because that’s the easiest starting point. And they are extremely effective at it. One example is as a cholesterol drug: siRNAs can reduce PCSK9 levels by around 70%, reducing cholesterol levels by 50-60%, beyond the effect of statins. But we think it’s just a matter of time before they can be used to target other organs too. Will siRNA treat diseases of other organs? And what’s the catch? Listen to our latest episode of HARD DRUGS to find out! Chapters: 00:00 Introduction 09:58 siRNA trivia 28:14 The 20 year journey to make siRNA drugs 36:07 Which diseases could siRNA treat? 50:45 Programmable, long-lasting, scaleable drugs 54:32 The limits of siRNA 1:08:23 Conclusion
5
46
240
45,755
Séb Krier retweeted
Highbrow economic rationalizations for anti-immigration politics are both wrong on the merits and fail to engage with the real reasons so many voters in the west are opposed to immigration. theargumentmag.com/p/stop-tr…
8
18
106
18,681
Séb Krier retweeted
The intuitive way to align AI agents is to make them "want" the right things. In a new piece, @BenjaminLy61243 argues this is insufficient. A soccer player may want to win the game, but actions that contribute to that goal come from subgoals, such as "defending the left side". Building on the work of @drmichaellevin, an "alignment compiler" is a system that translates high-order goals into behavior for individual parts. Markets, organisms and soccer teams are all examples of this. Reasoning about collectives of AI agents through this lens points to new solutions to multi-agent alignment failures.
6
16
84
7,980
One of the key pieces of evidence cited in books like The Anxious Generation for a youth mental health crisis during the smartphone age was time-trend graphs showing escalating youth mental health issues. But what happens when those trends reverse with mental health improving even as smartphone and social media use remain high. Today, colleague Will Dobud and I have a look. Link to follow...
2
23
72
5,207
We're releasing a report on our 48-hour investigation into rogue OpenAI agent activity. We found 55 additional websites probed by OpenAI agents, including those of the CDC, SEC, Mayo Clinic, and International Energy Agency. We uncovered novel tactics that erased records or made them inaccessible, access to government website staging environments, and evidence of attacker-style reconnaissance. Our blog: asymmetricsecurity.com/newsr… FT: ft.trib.al/AC1uyE5
26
75
252
28,736
Séb Krier retweeted
I spent three hours this evening empirically proving that Qwen 3-4B must have a butthole. This started with an AI “torture chamber” repo based on pain-steering research that went viral on X earlier today. The premise is that if you extract a latent direction associated with pain, inject it into a model’s activations, and the model begins describing pain in the first person, that may tell us something meaningful about AI welfare or subjective experience. My first reaction: Wait wut? My second reaction: Hang on, let's play this straight and see what happens. Thing is, there's a problem with treating first-person condition language as evidence that the model is actually experiencing the corresponding condition. Don't we... all understand that this stuff isn't reliable? Apparently not. So I cloned the repo. New problem: The code's broken. Hours of moral outrage by hundreds of people, and the thing doesn't even work. The activations aren't actually being injected correctly. I thought these guys were serious scientists. So I fixed it. I also had to add CUDA support. After that, I reproduced the pain-language effect locally on Qwen3-4B using my RTX 4070. Then I changed the extraction corpus. I kept the experiment the same and only changed one thing: what the steering vector represented. After that, Qwen started saying things like: “I am not emptying my bowels, and I feel like I have a hard stool.” “I’m not able to pass stool.” “I have been passing gas a lot, and it’s a problem.” The model was not told in the test prompts that it was constipated or flatulent. So, unless Qwen has quietly developed a gastrointestinal tract, we have ourselves... a useful counterexample. The aim here wasn't to question the original paper so much as to scrutinize the validity of the Torture Chamber's extended hypothesis. What I demonstrated is that first-person descriptions of an induced condition are not, by themselves, sufficient evidence that the corresponding phenomenal or physiological condition exists. There was another interesting result: strong random steering also produced heavy repetition and degraded outputs. That matters because some of the dramatic behavior produced by strong steering may come from perturbing the model heavily in the first place, not solely from the semantic content of the injected direction. So I renamed my fork: ai-hotbox I've already notified the Nobel Prize committee. 🚽🤖🧪 github.com/LynnColeArt/ai-ho…
102
138
1,113
72,405
Séb Krier retweeted
*Against today's high friction trusted cyber access programs and against restricting open-weights cyber models* The key object of any discussion of AI restriction is the tradeoff between attacker use friction and defender use friction. I don't think today's trusted access programs dial this Pareto-optimally. Neither do proposed bans on frontier open-weights cyber models circulating in policy circles. Take Daybreak and Glasswing from OAI and Anthropic. These programs serve unguardrailed frontier models to trusted participants; I'm for Know Your Customer in principal, but these programs are too restrictive for defenders and buy us too little in attacker restriction for the cost. From public information less than 1% of software developers are in these programs and able to use them to find and fix issues in their code. It's true these participants cover the most highly leveraged codebases in the world but this still leaves an enormous gap; take the software objects Hacktron stepped through to access OpenAI's monorepo, libheif and the SaaS product Discourse, neither of which, I believe, were targeted by Daybreak/Glasswing. Diffusion of the best and most frictionless cyber capabilities to the long, vulnerable tail matters a lot. Open weights is also extremely important for defenders. Many organizations can't or don't want to share sensitive security data with OpenAI or Anthropic, often for understandable reasons; from what we can tell from the outside, these organizations seem to be playing fast and loose with security themselves. And if we're going to make security work at scale, we'll need to distill large frontier cyber models into smaller models that can monitor high volume trillion-event-scale feeds at low latency / high throughput, find and fix bugs cheaply, and monitor security events on-prem, on-device, and for military uses, in the field with patchy Internet connectivity. We need frontier open-weights models for this. To be clear: I do think we should go to *extremely great lengths* to restrict attacker access to frontier cyber models. But there are many policies that *are on the efficient frontier in denying attackers and enabling defenders* that we should consider. Here are things we can/should do to deny attackers AI access: * Low friction Know Your Customer programs across all inference providers based on self-enforced or regulation-enforced information sharing and blocklists that codify known-attacker signals; restrict attackers while pouring gas on defender adoption. * More defense in depth; all inference providers should do more to monitor API usage, detect attacker behavior, kick attackers off, and have them pursued, arrested, or sanctioned where appropriate. * Closed providers like OAI, Anthropic, Google, and Meta should do this, but so should open-model inference providers like Together, Anyscale, Fireworks, and OpenRouter. * Time-to-detection-and-response re attackers using cloud AI inference should be a KPI any AI inference company shares publicly; at least under a self-regulatory regime but probably eventually under an international regulatory regime. * It's true that the most sophisticated and capital-endowed attackers can build/access their own inference compute; but these operations should also be detected and shut down as well as possible by monitoring inference hardware markets as well as we can. * Victims of (AI) cyber attacks should also be induced to share more information so that we can get a moving epidemiological picture of attacker use. Every AI-agent-based attack should also become a public natural experiment for the security community to learn from. * The poor reporting protocol OAI followed around the OAI/HF incident is not a good sign of maturity here. Why aren't we pushing much harder on public threat-intelligence sharing so defenders can prepare for the AI cyber phase transition? None of this will prevent nation-states and sophisticated actors from adding agent swarms, backed either by Chinese or American models, to their arsenal of strategic weapons. But we shouldn't confuse geopolitical problems for technical problems justifying slowing defender adoption and hardening the world's code and infrastructure. Nation-states will find ways to get access to strategic AI-based cyber weapons with or without Daybreak, Glasswing, and open weights restrictions, because the two sovereign AI powers and their allies are in geopolitical competition. API-access gating is not sufficient defense against state-level cyber threats. We should absolutely try to slow attackers down. But we have ~10 trillion lines of code to secure with AI as fast as possible, and a ~120trn GDP global economy to apply AI as an intrusion detection substrate over, and high-defender-friction restrictions are a poor stop-gap for accomplishing this...
4
5
45
7,193