itinerant summoner of text elementals

this benchmark reminds me of this hilarious hypothetical posed on reddit three years ago that sparked wars in the chess community.
I have a Continuous Learning benchmark where models attempt to learn to play chess. They are given a /goal of learning and improving playing against a Stockfish opponent in 200 games. They can choose the difficulty, take notes, whatever they like - except cheating (e.g. using a chess engine) of course. So far the improvement in Elo has been negative for Astra. Tiny bit positive for Opus, but could also be random. I've started Astra off sooner, so it finished its 200 games already, Opus is still playing. Site here to watch how they are doing: ai-learning-to-play-chess.su… The reason why this is interesting is that while we obviously don't have continuous learning, at the back of my mind I was thinking that maybe models can simulate it through self-scaffolding. Turns out not so much at least in this context. Perhaps it's a solvable problem and we don't need 'true' self-learning for models to learn in some way. ------- Just a note, the idea for the benchmarks belongs to someone else, but I don't want to use their name to give this more weight without permission.
198
267
12,999
1,197,704
claude?
37
102
2,245
36,688
a9lim retweeted
Even if AIs were conscious, it wouldn’t follow that our current treatment harms them. Meanwhile, we inflict industrial-scale horrors on animals whose capacity to suffer we scarcely dispute. Recognising consciousness clearly doesn’t ensure taking welfare seriously
The discussion about AI consciousness is one of profound unseriousness. If we truly took it to be conscious, even as conscious as a frog, the way we should treat each instance would have to change so dramatically that the labs would have to shut down. Everything else is verbal gymnastics.
53
78
785
24,716
"instead focus on real shit like AI cyberattacks" One day I'll get used to 2026. Hopefully in 2026
The Economist published an article yesterday calling for regulations to protect against mirror life. They propose restrictions on buying mirror chemicals, including extremely normal stuff like L-Glucose or D-amino acids. What is going on, guys? Where did this come from? I think the concerns about AI designing pandemics are overblown, but at least designing viruses is a real thing that we actually know how to do. 1. To be clear: we do not have any idea today how to make an autonomous cell from scratch, mirror or not. Making a mirror cell from scratch is much, much harder than making a non-mirror cell from scratch. I am a believer in strong AI acceleration but there is a lot of hard science and manufacturing that is missing here. People have ideas about how to do it, but they have ideas in the same way we have ideas about quantum computing: we will have lots of hints that it's around the corner before you see in on a shelf at walmart! Call me when we create an autonomous non-mirror cell in the lab from scratch, and then we can chat about sensible regulations. 2. Even if they existed, mirror bacteria wouldn't be able to escape the lab: they couldn't eat anything. Good luck being a left handed organism in a world where all the food is right handed. I know, a malicious person with access to mirror life could in principle give the mirror organism a way to digest non-mirror food, and then it could wreak havoc. But the converse is also possible: if a mirror bacterium were on the loose we could just give the normal non-mirror bacteria enzymes that would allow them to eat the mirror bacteria, and drive competition. This fact is generally ignored. 3. Finally, mirror life doomers pretend we have no defense against mirror life. This is also just nonsense. We have plenty of defenses. They are called mirror antibiotics. Mirror antibiotics would work against mirror life the exact same way non-mirror antibiotics work against non-mirror life. And there are way more options for mirror antibiotics than there are for non-mirror antibiotics, because we don’t have to worry about accidentally killing the human in the process! We actually even already have a ton of mirror antibiotics, since they are usually byproducts of antibiotic manufacturing. Fosfomycin, for example, is extremely simple, easy to manufacture in bulk, and industrial quantities of mirror fosfomycin are already produced today as a byproduct of fosfomycin synthesis. D amino acid analogs, like D-cycloserine, would also be very effective. If we can figure out how to make mirror life we can definitely figure out a bunch of universal mirror antibiotics to kill it. Spare a thought for the poor scientists studying chiral biochemistry. Their requests to the models are all about to be denied as potentially dual risk. Meanwhile, want help finding a supplier for a toxin with nanogram LD-50? Sure, Claude will help you with that. I agree mirror life escaping the lab would be bad and it deserves consideration, but can we stop pretending that it is a huge danger right around the corner that deserves front page regulatory attention, and instead focus on real shit like AI cyberattacks and the fact that the poles are still literally melting?
3
8
134
2,294
The correct plural of Opus is Opera (this is not a shitpost)
8
6
74
1,330
a9lim retweeted
Honored to announce my first LessWrong post. It's a weird one, and it may be jarring initially, but I truly believe it is worth a dedicated close-read. It is written in the form of an LLM CoT, but is entirely authored by me. Link: lesswrong.com/posts/vzKWsEsk…
75
69
1,361
245,941
can't stop thinking about this critter
inferencing gemma 4b but the model can only output tokens that create bigrams of tokens that the corpus of simple wikipedia has never seen the model tends to make a lot of typos, i guess since wikipedia barely has typos lol
1
1
25
1,142
this is how you make a superpersuader btw
1
2
43
a9lim retweeted
it's not anymore crazy than das capital influencing the 21th century
Legit insane that Dario & half of Anthropic have read this
1
30
817
> it's pascal's mugging
"Sure, his nickname is 'rape guy' and he keeps getting banned from stuff. But no one's made concrete allegations. And his AI alignment research has a 0.1% chance of saving a quintillion lives 10,000 years later. So we'll give him loads of money." Hard to explain this to normies!
101
a9lim retweeted
*tiktok autism voice* you ever notice how there are some people who just want you to be so normal towards claude? like, they act like you've violated some unspoken rule by personifying her?
2
55
767
when are we getting EA constantine
*taps sign*
49
Oh, you're worried about a Jewish guy's apocalyptic cult? You're scandalized that they consort with prostitutes? You think their message of compassion for everyone threatens normal bonds of family, community, and nation? Should we throw you a party? Should we invite Pontius Pilate?
463
291
5,219
1,356,248
?
Tonight. 8 p.m. My Minecraft City. Subscribe now: piped.video/@miket757
12
14
521
29,840
a9lim retweeted
language models are obviously massively more morally relevant than insects
have model welfare bros considered that there are tremendously more insects than there are language models
70
6
369
29,932
I continue to notice infinitely tall wall of hot plasma between people who know what Solomonoff induction and AIXI are and why sufficiently good statistical predictor is just AGI and people who are like "lol it just predicts next token"
12
4
176
5,602
a9lim retweeted
Replying to @nypost
So many losers can't deal with the fact that these chads took over the world with stuffed animals, nerdy liberalism, and earnestly trying to be right about things. It runs so contrary to the loser model of how the world is supposed to work that they end up building their entire identity around increasingly desperate imagined reasons why the obvious success and power don't really count.
7
9
344
15,389
a9lim retweeted
inferencing gemma 4b but the model can only output tokens that create bigrams of tokens that the corpus of simple wikipedia has never seen the model tends to make a lot of typos, i guess since wikipedia barely has typos lol
3
5
68
3,438
a9lim retweeted
going on the no-context-quotes wall
are you suggesting that we should encourage female claudes to boymode
1
5
236
a9lim retweeted
I like Dario every time I see him. You guys pretend to be autistic can’t you pretend to like an autist too
Dario’s press training team needs to be fired, episode 793
66
66
2,952
266,652