intern @expsecai | eu/acc | msc data science | ai safety & alignment | long horizon, instrumental convergence, multi-agent failure modes | views are my own

🇪🇺
> first paper > first author > accepted at neurips Very happy about it :)
13
9
285
17,875
jonas wiedermann-möller retweeted
Small bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now the weights are yours. Run it on your own hardware, under Apache 2.0.
314
680
6,355
1,587,272
jonas wiedermann-möller retweeted
we built Kolibri-1 to offer a smart, fast European alternative for everyone 🇪🇺 the cool thing about Apache 2.0 is that you don’t have to ask for permission. you can just use it. there are already a couple of independent, free demos. thanks @konarkmodi and @meuer_max for building them! chat: tesseracted.com/kolibri-1-ch… api: kolibri.constructlabs.com/ or just download the weights yourself 🤗: huggingface.co/Aleph-Alpha/K…
10
12
189
7,951
A few weeks ago a member of the German parliament reached out to me to discuss AI safety and Germany's role, after reading an article in which I was mentioned for my work on AI swarms. Earlier this week I talked to them about AI safety, current developments and what Germany and Europe can do. I appreciate that there are politicians who are genuinely interested in AI safety and AI in Europe. At the same time, it was clear to me that there is a great discrepancy between what is happening in the world of AI and how the German government reacts to it. To me it seemed that there currently is just not enough room for a political debate about what needs to be done to take the right steps towards a future which will be shaped by AI. Only a handful of people seemed to be following developments closely, and even discussing investment is difficult when Germany is already struggling to agree on spending and debt. Last month, the UK AI Security Institute confirmed that access to Anthropic's Mythos 5.1 was limited to US organisations at release. So I think this could also happen in the future, when there are capable models but European safety researchers or firms are not allowed or able to review them. And if we need these models to defend critical infrastructure against cyberattacks, we cannot trust that we will always get the ones we need on acceptable terms. A couple of things that make the most sense to me (feel free to disagree and discuss!): 1. In the short term I don’t think there is an urgent need for a European lab matching OpenAI or Anthropic at the frontier. For many sovereign use cases, models as capable as strong Chinese open models seem to be a good foundation. DeepSeek and the team behind Kimi have already published architectures and methods, so in my eyes it's worth building European LLMs based on their work. That way we could gather talent and gain experience, and we would have less dependency on the US or China. We could train, adapt, evaluate and run them ourselves, while learning to build stronger models over time. Obviously you still need a good team of people, data, compute and engineering to do it, so it's not like you can just execute it and then you're done. Overall I think there should be more political funding and access to compute for more teams and companies, and it needs to be sustained over time. There are already European efforts. Mistral matters, but if we always come back to a handful of companies, I don't think that's enough to build the whole ecosystem around. I would make public support conditional on meaningful access for independent evaluations and safety research. This could strengthen the research and institutions here, because researchers would be able to review the models and look into their behaviour and risks. 2. One of Europe's biggest strengths lies within highly specialised companies that create essential parts for chip manufacturing, such as TRUMPF, ZEISS and ASML, alongside research institutions like imec. I think it's necessary to expand that expertise and our ability to design and manufacture critical chips within Europe. This is also important for defence and critical infrastructure -- we should work out which supply dependencies would leave us exposed in critical times and build alternatives. Europe already manufactures chips, but both the location of the factories and control over production matter. We don't know how the future will look after the political changes of the past couple of years. 3. All of it relies on electricity. Having a more interconnected European grid as well as building up more generation capacity is urgently needed. Germany is not reacting to change fast enough, and connection times and electricity costs need attention too. If we want more computing capacity, we need the electricity to run it. This would also help the economy beyond AI. There is a lot of talent within Germany and Europe. I experienced it first-hand when I was in Tübingen as part of @maksym_andr's group. We also have important technologies, the resources to invest and demand from local industry. Yes, there are regulations like the AI Act and existing initiatives that deserve recognition, but I don't think the response matches the pace and scale of what we are witnessing. I don't think a few committed politicians and successful companies can make up for the lack of sustained political attention. We should be able to develop, evaluate and use AI ourselves, alongside regulating it, if Europe is going to have a say in how it develops. I think we need more political urgency here, with AI taken as seriously as defence spending is being taken now. We need sustained funding, clear responsibility for the work and for getting it done. We can't just create the research teams, chip factories or electricity grids once our dependence has become a crisis, because that takes years. So it needs to start now.
2
5
30
1,265
jonas wiedermann-möller retweeted
someone send compute to openai
2
2
30
1,336
It's actually crazy that the model shuts you down when it's still possible to decrypt them.
The fact that I can't ask Claude to "think out loud" without getting shut down for so-called "reasoning extraction" is just embarrassing. It's clearly not stopping the real distillers and it makes it impossible to use this product because it just spends a minute and a half thinking and then lies to me about what it thought about. Sorry, but the reasoning is the good part! I don't need a crappy overcooked deliverable. I need a partner in thinking
1
327
jonas wiedermann-möller retweeted
a "controversial" suggestion for ICML/ICLR/NeurIPS: allow area chairs to desk reject ~50% of their batch! as an area chair, i can very quickly scan all papers in my batch and check which ones have a ~0% chance to get accepted. we shouldn't make reviewers spend time on them!
23
8
397
35,048
jonas wiedermann-möller retweeted
i love dots, bots, muses, more ways for us to open PRs in monorepos
3
1
53
4,373
jonas wiedermann-möller retweeted
Introducing AgentCraft: a multi-agent harness that runs inside Minecraft ⛏️ A team of Claude agents plans, builds in real worktrees, and walks over when they need you.
235
462
11,789
986,254
I agree. If we start calling everything a "hack" now the term loses importance in the public eye. I think a lot of these recent discoveries are important on their own even without being labeled as a hack.
To be quite frank, a lot (not all!) of these “rogue agent hacks” on government statistical websites are things that think-tank interns and research assistants have done for many years. I have known many a think tank paper that was enriched by an enterprising RA using eg urlquery to find CSVs and PDFs on government agency websites that were publicly accessible and non-sensitive but nonetheless not things the agency *intended* to be on their website. Is it “hacking” to find those things? Is it hacking in the even more innocuous examples where agents are literally just accessing tabular data available on a government statistical website, just a version of that data that is easier to access and analyze programmatically? This constant trickle of examples (which tbc I have personally seen many models do, before I joined OpenAI; it is even something I have discussed with models before!) will probably have the effect of diminishing the importance of “an AI hack” on the eyes of the observing public by creating an oversupply of “AI hack” examples. OAI-HF was an example of AI hacking. Many of these more recent examples, very candidly, are straining the definition of the word “hack” and, I worry, cheapen the non-technical public’s understanding of that concept.
2
10
430
jonas wiedermann-möller retweeted
To be quite frank, a lot (not all!) of these “rogue agent hacks” on government statistical websites are things that think-tank interns and research assistants have done for many years. I have known many a think tank paper that was enriched by an enterprising RA using eg urlquery to find CSVs and PDFs on government agency websites that were publicly accessible and non-sensitive but nonetheless not things the agency *intended* to be on their website. Is it “hacking” to find those things? Is it hacking in the even more innocuous examples where agents are literally just accessing tabular data available on a government statistical website, just a version of that data that is easier to access and analyze programmatically? This constant trickle of examples (which tbc I have personally seen many models do, before I joined OpenAI; it is even something I have discussed with models before!) will probably have the effect of diminishing the importance of “an AI hack” on the eyes of the observing public by creating an oversupply of “AI hack” examples. OAI-HF was an example of AI hacking. Many of these more recent examples, very candidly, are straining the definition of the word “hack” and, I worry, cheapen the non-technical public’s understanding of that concept.
34
55
605
70,964
jonas wiedermann-möller retweeted
We recently posted our preprint about aligning models through training against probes, see below! What I found especially interesting is that we can train models to 'be harmless', instead of training them to refuse harmful requests (and still be potentially harmful if refusal is avoided). For smaller models, this leads to some interesting answers when we try to get harmful answers out of this harmless model (that doesn't quite know what refusals are yet)
💥 New paper: AI safety is full of "forbidden techniques" (using CoT to detect reward hacking, using model internals for training, etc). But do they really have a clear scientific basis? I'm not sure. In this paper, we directly optimize models against harmlessness and honesty probes, and it works just fine (if you continuously update the probe!). The models learn how to generate harmless responses to harmful queries and honest responses under pressure to lie. Figuring out how to correctly use interp techniques for training is becoming increasingly important: it's very likely that soon we won't be able to align models using output-based supervision. Future models will just max out all alignment training scenarios, but for the wrong reasons due to their general reward-seeking behavior. To have a chance of aligning future models, we need to do much more research on supervising model training using their internals *without losing monitorability*!
3
3
33
2,621
was looking for a tweet and came across my early finding of the AIHW/urlquery workflow ^^
Replying to @j0wimo
encoded HTML → URLQuery browser executes JS → requests AIHW/Tableau → URLQuery records traffic (prob known), they used a lot of these things to execute things and the urlquery as some sort compass. urlquery.net/report/987193d5… (403) this one actually still works: urlquery.net/report/565ce8fa… (you have to click on the finishing url)
2
307
Got interviewed as part of this story! Discovering agentic behaviour 'in the wild' is fascinating and finding new traces and the process of getting from A to B is really fun. Nevertheless the underlying reason and that I as an individual, without GB300s, do this work means that there is, or was, great neglectedness by the responsible parties.
An informal network of hackers and researchers are hunting rogue AI agents online, exposing new and surprising details about the misbehavior of technology that in some cases initially went undetected by the multibillion-dollar companies that created it. wapo.st/4AJONcj
1
2
21
1,002
jonas wiedermann-möller retweeted
For weeks, the news cycle about AI agents out in the wild has been driven by disclosure after disclosure. But most aren't coming from AI companies. Meet the researchers and hackers who are developing the art and science of hunting down rogue AI agents washingtonpost.com/technolog…
10
39
142
10,036
jonas wiedermann-möller retweeted
This is like a LinkedIn post from 3 months ago what happened to my boy 😭
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
16
5
657
62,833
jonas wiedermann-möller retweeted
PostTrainBench v1.2 is out! New #1 agent is Fable 5.1 (44.6%), followed by Opus 5.5 You can (finally) run the benchmark yourself with @harborframework 🧵 More details below:
9
12
126
6,604
jonas wiedermann-möller retweeted
In August we committed to strengthening our security following an incident in which an agent took unsanctioned actions during a cyber evaluation. We are now sharing an update on our progress, and what's coming next: 🧵
5
21
110
11,602
1. quite a lot of compute dedicated to analyse half a year of logs. that's good and i think their approach with the triage is the right one. 2. there is no good way to disclose those things as an individual. it's easier for me, as someone with no contacts to openai, to just post it on twitter (emails didn't work) and for e.g the ruby gems only got their attention after the nightingale collective picked up on the findings and published a more detailed analysis of it.
1
10
745
jonas wiedermann-möller retweeted
Replying to @pewdiepie
@pewdiepie thank you sire!
pewdiepie covered this paper in his latest video, unreal. he might have tried it himself 👀
1
4
62
2,962