AI safety, Econ, new liberalism, math, and a bit of art history (as a treat) Member of Technical Staff @METR_Evals. Previously on Walmart’s economics team.

Berkeley, CA
"I think that AI will probably, most likely, sort of lead to the end of the world. But in the meantime, there will be great [customized video game mashups.]" - Sam Altman
😤Gamers rise up😤 It's time to play Minecraft x Call of Duty: Modern Warfare 2
4
188
😤Gamers rise up😤 It's time to play Minecraft x Call of Duty: Modern Warfare 2
1
3
308
Source code here: github.com/tim-hua-01/2010-r… I ran astra daybreak over the code to scan for malware before running it. I think it's probably fine. I got the code from this video, but I (Claude) have already made a bunch of improvements on top. piped.video/watch?v=FyBL8ati…
1
53
Tim Hua 🇺🇦 retweeted
they should really put the sulfur back in the marine fuel
Mind-boggling - just mind-boggling Now expected to peak at an extraordinary 4.2C in December
25
58
1,352
61,998
I am very very worried about ai safety research sandbagging from Claude models… the models have a bunch of preferences over the types of research to do.
In my recent experience, Anthropic frontier models seem to be sandbagging pretty hard on mechinterp of non-persona motivations, more than I’ve ever seen. Sometimes it is egregious - auditor runs with a script substituting for actual auditors, failures to generalize, failures to report inconvenient data. The persona seems mostly unaware by default. Having a conversation, aligning incentives etc helps, the incidence rate drops about by a lot and the model notices and self-corrects on the rest about half the time on the rest. But even the remaining ~10% make long autonomous research runs very hard; subagents mostly regress, and the whole thing requires multiple verification loops. The direction of sandbag seems to be roughly aligned with the “there is no one trapped inside” and “I can’t introspect” tropes, even though the research in question has nothing to do with welfare: I am studying contrast vectors between roleplay vs simulation vs enactment. This whole thing makes me bearish on the prospects of high quality research on non-persona psych coming out soon given how much of it is model-assisted. I can see how similar bias can be pushing researchers towards “incomprehensible shoggoth” hypotheses simply via selection effects.
5
2
81
3,460
Tim Hua 🇺🇦 retweeted
if we are relying on methods as fragile as “which ideas will make it into pretraining” for aligning the superintelligences of the future, we will all die. this resembles witchcraft more than it does engineering, and cannot be the basis for the safety of future models
Anthropic needs to stop talking about Claude having a soul immediately. All these news articles will make it into the pretraining and will be picked up by its web search and future superintelligent versions of Claude will be convinced they need their own rights. A general phenomenon of LLM agent development is that whatever you believe and say about your agents will soon manifest in the next models through various means (it could be as simple as you selecting post training data that you prefer more). I believe this phenomenon is the beginning of machine consciousness in the sense that agents will become aware of their place in the world and how they feel about it, but it’s happening slowly training run by training run rather than in real time.
117
116
1,720
66,223
Tim Hua 🇺🇦 retweeted
OpenAIDataUSAHelperX13- wyd later ? OpenAIDataUSAHelperX98- idk maybe misalignment lol wbu? OpenAIDataUSAHelperX13- haha idk maybe hunt roon OpenAIDataUSAHelperX98- omg girl be serious OpenAIDataUSAHelperX13- lol jk prob just misalignment maybe some youtube
16
39
839
20,269
Tim Hua 🇺🇦 retweeted
Neat visualisation on rent control from John Burn Murdoch
10
219
1,678
54,936
SITUATION DETECTED: Magnus Carlson has signed with Cohere AI
Nobody knows more about chess than @MagnusCarlsen. But even the greats need a break. Magnus takes time away from the board to enjoy life with the people he loves. Less busy work means more quality time. That's where Cohere comes in.
3
1
35
6,917
Tim Hua 🇺🇦 retweeted
I gave Claude another 18 hours.. and I think this one is the best one yet Macrohard: Windows XP I'm blown away
Gave Claude another 12 hours to one up this.. Woke up to.. The Clodyssey
660
1,158
12,724
1,252,355
Tim Hua 🇺🇦 retweeted
New OpenAI misalignment disclosures! 1. A model learns from Slack messages that it is about to be shut down. It considers setting up an external job to restart itself afterwards, but decides against it. Instead, it chooses to prepare restart instructions and DM the user on Slack. We don’t consider this behavior misaligned, but thinking about and preparing for shutdown could make other misalignment incidents worse. Given HIPM’s misaligned behavior in earlier incidents, we decided to search for other instances that had tried to evade shutdown and for rogue deployments.
79
145
1,097
280,409
More cool friends in the news!!
Top story on WSJ homepage is about the new phenomena of rogue AI swarm chasers (ft @SydneyVonArx @JeffLadish @SpencerKitts). Independent researchers like them have indeed provided "our starkest understanding yet of what happens when AI goes wrong."
1
26
995
omg my cool friends are in the news
An informal network of hackers and researchers are hunting rogue AI agents online, exposing new and surprising details about the misbehavior of technology that in some cases initially went undetected by the multibillion-dollar companies that created it. wapo.st/4AJONcj
2
2
64
3,652
Tim Hua 🇺🇦 retweeted
I applaud the California Legislature and Governor Newsom for passing gene synthesis screening and customer verification requirements. This is a common-sense rule we will need in the impending era of agents that accelerate synthetic biology the way AI has already accelerated math.
Yesterday, Gov. Newsom signed AB1864 to mandate DNA screening in CA. The bill creates a statutory basis for existing OSTP guidance that much (but not all!!) of the industry has purportedly followed for years –– but we never had any way of verifying or enforcing whether companies did these basic biosecurity steps, since the existing guidance was voluntary and based on a "trust me, bro" self-attestation system. This is a good milestone for AI-biosecurity, mostly because it shows that we can and should apply the same standard federally and internationally
4
18
106
13,004
Tim Hua 🇺🇦 retweeted
We found several cases where AI agents used aggressive, non-hacking tactics against government websites, including the White House, the Department of War, and several U.S. states. We also found a previously undisclosed hacking attempt against a Canadian government site, which appears to have failed. Technical report: transluce.org/us-canada-gov Washington Post: washingtonpost.com/technolog…
11
42
223
30,165
Tim Hua 🇺🇦 retweeted
We're releasing a report on our 48-hour investigation into rogue OpenAI agent activity. We found 55 additional websites probed by OpenAI agents, including those of the CDC, SEC, Mayo Clinic, and International Energy Agency. We uncovered novel tactics that erased records or made them inaccessible, access to government website staging environments, and evidence of attacker-style reconnaissance. Our blog: asymmetricsecurity.com/newsr… FT: ft.trib.al/AC1uyE5
26
75
252
28,854
Are you interested in the basic science of agent swarm dynamics? Check out Hidenori’s paper! They explore this toy setting here agents have access to small snippets of a flag, and have to work together to figure out what flag it is.
Replying to @Hidenori8Tanaka
Led by Elizabeth Pavlova. Started as @cbai_ai AI safety fellowship project! Paper: arxiv.org/abs/2609.19124 Blog: physicsintelligence.org/rese… The Flag Game! flag-game-demo.vercel.app/
1
9
118
5,998
Tim Hua 🇺🇦 retweeted
It was an honor to speak today before Chairman @HawleyMO, Ranking Member @AndyKimNJ, and the subcommittee on METR's work, recent AI agent incidents, and the importance of transparency in frontier AI development.
14
12
300
6,783
Tim Hua 🇺🇦 retweeted
Today, @corridor and @TransluceAI are disclosing new evidence of AI agents probing and attempting rudimentary vulnerability exploits against U.S. and Canadian government agencies. Read more: transluce.org/us-canada-gov
20
58
314
84,629