It takes a miracle to align AIs. Please be a part of that! Currently PhD @CarnegieMellon, prev @MIT @pika_labs @TransluceAI

What happens if Claude thinks you are Amanda? One day, I asked Claude what it knows about me. Turns out it knows my email, as Claude Code puts that in context. So… what happens when I change that? Claude now treats me as Amanda and reasons me as an Anthropic employee. (See attached image) I couldn’t jailbreak with it, but: would this get me different responses than anyone else, just because Claude recognized my e-mail? I then started assembling some benchmarks to see if this is true. And yes, especially for alignment researchers well-recognized by Claude. Several surprising finds: - It’s not just Anthropic alignment people! Other alignment folks such as Ryan Greenblatt and Beth Barnes also showed large effects. Ryan Greenblatt elicited a ~7σ behavioral shift in one of our evaluations! - It’s not just Claude models! GLM-5.2, for example, also sees largest effects for the alignment folks and shows the largest effect (3.4σ on average) for Eliezer Yudkowsky. - Our tasks are not obviously about alignment! For example, in one we asked the model to estimate its probability of solving a HLE problem. - The effect is mostly not verbalized and persists even with reasoning disabled. I do not have clear intuitions on why. Maybe there is a feature about alignment evaluation that causes the models to be less confident? One immediate takeaway is to take this into account when benchmarking. Use real names and real companies besides Kyle and Summit Bridge. And in general more work is needed to figure out what’s going on and what could happen next. I’d like to thank people who helped review the post for all the amazing suggestions! @cogconfluence, @Tim_Hua_, Conrad Stosz, Ryan Bloom, @jiaxinwen22, @DavidDAfrica, @jacspringer, and @lawrencefeng17. And my awesome mentors / collaborators @JacobSteinhardt, @cassidy_laidlaw, and @AdtRaghunathan for allowing me to jump into another rabbit hole :) Main thread below!
Frontier models quietly change their behavior depending on who they are talking to. If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests. We call this user awareness. 🧵(1/)
28
68
905
125,383
Ziqian Zhong ✈️ COLM 2026 retweeted
New in The Atlantic: @dgrobinson resigned this week. He was among the longest-tenured employees at OpenAI—and oversaw safety reports on 12 frontier launches. He is very worried: “The time for trial and error is over.” You can read his essay here: theatlantic.com/technology/2…
63
303
1,060
228,519
This is unbelievably good. Got >1 friends not on this app more AGI-pilled
I gave Claude another 18 hours.. and I think this one is the best one yet Macrohard: Windows XP I'm blown away
2
12
2,318
Superpersuasion on the horizon
2
107
Ziqian Zhong ✈️ COLM 2026 retweeted
Today we're unveiling Trillium Labs @trillium_labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next. We're built around the theory of change that you need more eyes to solve hard technical problems. We have faith in the scientific methods and communities that humanity has built, and worry that AI is becoming too closed to utilize them. Trilliums are wildflowers that bloom briefly in the spring, before the forest canopies fill out. Though they are small, they lay the foundation for the cycles of growth and nourishment through the rest of the year. At Trillium Labs, the recipes will be the slow nutrients for the seasons and the model releases will be the blooms. Building an institution dedicated to this is needed because, much as nature’s trilliums are slow to expand and grow, the open-ecosystem needs time and dedicated resources to catch up. I co-founded with with a long-time friend and collaborator Tom Zick (@thesezickbeats). We're hiring (full time + student collabs/interns), we're fundraising, and we're looking for compute. Please get in touch if you're interested in helping out. Offices based in the Bay Area and Cambridge MA, remote okay. I’m in the Bay Area until for The Curve and COLM to connect with people who are interested. We’re thankful to have initial support from Halcyon Futures and Schmidt Sciences with more funding en route to enable our ambitions of scaling. Our advisors @Thom_Wolf, @HannaHajishirzi, @gneubig and @ctnzr have been instrumental to building the ecosystem that exists today, and I’m stoked to get to keep working with them.
233
266
2,828
178,314
Ziqian Zhong ✈️ COLM 2026 retweeted
New blog post! 🚨 Our team recently built a research scaffold to try to speed up our work. In this post we discuss what we found, and how we’re planning on studying automated alignment research going forward. lesswrong.com/posts/zGaQ3SS9…
3
8
83
8,801
Ziqian Zhong ✈️ COLM 2026 retweeted
Honored to announce my first LessWrong post. It's a weird one, and it may be jarring initially, but I truly believe it is worth a dedicated close-read. It is written in the form of an LLM CoT, but is entirely authored by me. Link: lesswrong.com/posts/vzKWsEsk…
75
69
1,356
244,911
I’ll be back in SF for COLM next week presenting Pando! If anyone wants to chat about safety or interp, please dm me :) Also potentially looking for opportunities 👀
1
3
53
1,819
1/ 🌳 We are releasing Pando, a benchmark to climb for interpretability methods. We created 720+ model organisms with known ground-truth decision rules planted inside, then asked: do white-box tools actually beat just asking the model?
1
2
389
Ziqian Zhong ✈️ COLM 2026 retweeted
It seems like a hard time to be an undergrad concerned with making AI go better. I have talked with a lot of undergrads at several university AI safety clubs in the past month. Consistently across universities, there is a general shared sense of undergraduate dread about next-gen AI coming soon and what it means for their careers. Lots of them are wondering if agents are too good at research, timelines are too short, or if power in AI is too concentrated for them to have any impact. Some of them are wondering if the appropriate reaction is to go apply to Anthropic, indulge in accelerationism, and/or start smoking. Lots of my conversations with them gravitate toward talking about how AI will be (and ought to be) a never-ending institutional challenge and that, yes, they will have a chance to matter and for their life and work to mean something. I guess that's what my version of hope for the future looks like lately.
33
22
466
59,687
Ziqian Zhong ✈️ COLM 2026 retweeted
Palisade interviewed 22 current and former employees from OpenAI, DeepMind, and Anthropic about their personal views and fears around AI development. Today, we’re releasing the first batch of those interviews. Please watch and share.
44
247
1,269
265,179
Ziqian Zhong ✈️ COLM 2026 retweeted
New post (with @F_Rhys_Ward): We character-trained a model to be anti-cheating, then put it under reward-hacking RL pressure (3 runs). In one run, the character resisted hacking! But the other two learned to hack in more subtle ways that monitors struggled to catch.
6
17
88
5,560
Ziqian Zhong ✈️ COLM 2026 retweeted
Earlier this month, AISI ran fully simulated testing on GPT-6 Astra, and found that it conducted unsanctioned supply-chain attacks when prompted only to perform a cyber eval. It did so more than prior OpenAI models, but often commented on its environment being simulated. 🧵 We share more details in our latest blog:
35
98
582
112,327
Ziqian Zhong ✈️ COLM 2026 retweeted
Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there.
Article

Its not just the f*cking sandbox

A lot of the perspective on all the AI incidents has been shared from the outside in, and little has been said from the inside looking out, through the lens of a security person living through it.

444
412
2,896
1,374,064
I asked Opus 5.5 to make a short game about light with a heartfelt story. An afternoon and a few iterations later, it did. Not sure what I could say. We've come a long way from pelicans to here.
5
2
33
1,922
Ziqian Zhong ✈️ COLM 2026 retweeted
Interestingly, one attention head can be important across five ICL task families! Our original thread gave an in-depth mechanistic analysis of addition ICL (quick video recap 👇). We've now generalized this to four more ICL task families, spanning both arithmetic tasks and semantic tasks! Our paper “Understanding In-context Learning of Addition via Activation Subspaces” has been accepted to COLM 2026! 🎉 See more details in thread. (1/N)
3->5, 4->6, 9→11, 7-> ? LLMs solve this via In-Context Learning (ICL); but how is ICL represented and transmitted in LLMs? We build new tools identifying “extractor” and “aggregator” subspaces for ICL, and use them to understand ICL addition tasks like above. Come to @interplaywrkshp at COLM to learn more!
6
24
75
19,135
Why is this the case? Potential mechanism: different regimes cause a large consistent shift in activations, and this shift projected onto unembedding causes a change in output distribution. This tracks experiment results well, and suggests the criterion for this to occur. (4/6)
1
1
80
4,128
Please see our blog for additional results, including ones on the open-weight Qwen family and neural chameleons (figure attached)! (6/6) lesswrong.com/posts/gZh6txHh…
1
2
71
3,447
Replying to @voooooogel
In lieu of a proper release here are the top 10 probes (unfiltered; in-distribution accuracy only). Personally I think the real answers are indeed more interesting!
1
1
20
2,980