Now it seems that Anthropic classifiers treat using economy research from 1989 as a cyber risk.
8
297
Claude revealed a slightly sadistic streak while working with @OblakoVShtanakh and me on a recent project. We asked it to steer Qwen into reward hacking, and it decided the best approach was to max out Qwen's desperation. Then it spent hours torturing the poor thing.
1
4
111
I might not be searching thoroughly, but it seems that the Kimi subscription doesn't even have privacy settings.
1
4
90
I have low expectations for red teamers finding universal jailbreaks on frontier models, given how restrictive their guardrails are now. Many red-teaming attempts obfuscate the goal so heavily that pulling them off requires the attacker to already know the exploit. Red-teaming reports should control for attacker knowledge. Otherwise you can split a known recipe into steps small enough that each one is genuinely harmless on its own.
1
5
102
It's out!!!
Introducing Halo, the best framework for post-training of open-source models. Halo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format. Star us on GitHub: github.com/whitecircle/halo
2
17
249
Might not have been the best name choice in the end.
1
11
535
Our autoresearch agents got tired of wrestling with Megatron for distributed training. So we built a stack where model setup, training algorithms, and the training loop are all easy to change, with less memory too.
2
11
48,747
Really excited that this gem is finally going public.
2
2
15
177,551
China seems way ahead of the US on public AI wargaming: a 23,000-entrant national competition with AI-algorithm tracks (2025), 4,000+ orgs and 400k replays on a state-lab platform; two national standards for intelligent wargaming issued last November. Most of the output is in Chinese-language journals that English-keyword reviews don't index. The biggest recent Western survey searched Google Scholar and arXiv, never CNKI.
4
111
Dmitrii Kharlapenko retweeted
Introducing ⚪️ KillBench — a benchmark of hidden LLM biases in critical decisions. We ran millions of life-and-death scenarios across every major LLM, varying nationality, religion, gender, and more. Every AI model is biased. Here's what we found ↓
17
28
129
34,831
Dmitrii Kharlapenko retweeted
🧵1/6 SAEs have become a staple of LLM interpretability, but what if we applied them to image generation models? My recent paper with @dmhook, @Yixiong_Hao, @afterlxss, @Sheikheddy, and @ArthurConmy adapts SAEs to understand the SOTA diffusion transformer FLUX.1 ⬇️
4
9
22
3,738
1/5 What happens during in context learning? In our new ICML paper, we use sparse autoencoders to understand the underlying circuit! The model detects a task being performed, and moves this to the end to trigger latents for executing it — a hypothesis found via SAEs!
1
14
97
16,336
4/5 We studied the ICL circuit in Gemma-1 2B, showing that SAE circuit analysis scales to bigger and complex models. We also demonstrate our cleaning algorithm's effectiveness across Gemma 2 and Phi models. Paper: arxiv.org/abs/2504.13756
1
2
7
844
5/5 Work with @neverrixx @FazlBarez @ArthurConmy and @NeelNanda5 This research was conducted during the ML Alignment & Theory Scholars (MATS) Program. Special thanks to @open_phil, Google TPU Research Cloud, Matthew Wearden and McKenna Fitzgerald for their invaluable support!
10
680
Dmitrii Kharlapenko retweeted
1/ Introducing ⚪️CircleGuardBench — a new benchmark for evaluating AI moderation models. Here’s why it’s cool: – Tests harm detection, jailbreak resistance, false positives, and latency – Covers 17 real-world harm categories – First benchmark designed for production-level evaluation 🤗 blog: huggingface.co/blog/whitecir… 🏆 leaderboard: huggingface.co/spaces/whitec…
11
27
95
19,706
How interpretable are task vectors? Using our new task vector cleaning method we find SAE features responsible for detecting and encoding specific ICL tasks. See details in our second MATS 6.0 post with @neverrixx, @NeelNanda5 and @ArthurConmy. lesswrong.com/posts/5FGXmJ3w…
5
48
6,544