Built @RossumAi, AlphaGo baseline pachi, bits of git, elinks & other oss... "The world is awful. The world is much better. The world can be much better."

Prague
Pinned Tweet
Summarizing my most important tweets from the last two years on consciousness, future, and the narrow winding path through the singularity: pasky.or.cz/singularity/ AI and alignment are the final exam of philosophy.
1
3
544
Petr Baudis retweeted
Replying to @Noahpinion
Took the time to digest @MiTiBennett's paper on consciousness smeared over time, and I'm a bit confused how it is relevant in practice: 1. There always seems to be some fine enough time granularity on which some ingredients don't co-occur. Trivial example: did you know that LED matrix displays always show only one column at a time? they just switch columns very fast. 2. How this is solved is through integration elements that *do* smear them over time. They can have a variety of forms. 3. There seems to be no reason why wouldn't we have machine perception integration with the same properties to physical ones. 4. The integration elements can be cognitive too. E.g. lots of biological vision processing gets trained directly on the retina neurons. 5. In extreme cases, you could have time steps in LLMs not with token granularity, but with turn or even context window granularity. Suddenly, you get concurrency it seems much higher than the human brain can do? 6. Even if you stick to the token level and give some magical significance to the hardware layer, attention is massively polycomputational. Much of the cognition / consciousness realization happens when computing the very first token of the response, and basically all the pre&mid training goes into this getting more and more powerful. You still get limited by as much as O(|context window|) ingredients - which seems massive? Did I miss something obvious?
1
1
151
Petr Baudis retweeted
Philosophy of mind discourse on Twitter is the specific place I consistently see otherwise very smart people say the equivalent of "Well moron rocks could never orbit the Earth because their telos is to be close to the ground, everyone with half a brain knows that."
19
28
368
19,192
too fucking cyberpunk
虽然知道这是为了給具身智能采集数据赚点外快补贴家用 但夜市看见这个还是太他妈赛博朋克了
1
3
639
my thoughts are a bit different: how good do we think IBKR's security really is? I'm genuinely unsure how to hold stocks securely during the potential cybersec apocalypse. (Banks are different, you can stop / undo wires, but trades?)
Fascinating. Chief Economist at Apollo: agents could cause a bank run by sweeping household cash into accounts paying 3-5% instead of the 0.1% national average, causing banks to lose a large share of their cheap deposits.
4
7
1,702
more generally: you really need to *immerse* yourself to develop taste we used to have flow states and thinking about and playing with an architectural problem for days find other ways to immerse to develop the types of taste that still matter
*deep breath* people think I'm joking when I tell them that to develop taste for software they need to have brunch with their friends, watch old movies and find old records, walk around metropolitan cities and observe how the building styles lend to the personality of the people around them, but I'm so fkin dead serious. where else is this mythical "taste" going to come from? what are you doing to parse the human condition? mf you can barely decide what discomfort you're subjecting yourself to in finding out what you like/dislike, and I'm supposed to trust you to make decisions for millions of people? go do some pottery and come back to me and describe the joy you felt without faffing about curves or whatever, I need you to feel the clay under your fingernails and fall in love with mud. even steve had the decency to drop acid on a sweaty beach in goa to find himself, the least you can do is listen to your friends talk about their lives over a meal without getting distracted by your phone. *exhale*
1
5
1,198
Petr Baudis retweeted
The paradox is that once a model is able to oneshot something, it becomes worthless. Doesn't matter how cool it is, since everyone will produce exactly the same and nobody wants to consume it.
Opus 5.5 on Max effort - "make a dynamic 15-second motion graphics video that shows what an incredible motion designer you are, like it's your showreel for a résumé. go all out."
418
250
5,922
396,885
> and what technologies could do the same today?
I had not realised before that the Bronze Age Collapse was literally *caused* by iron. Bronze age powers like the Egyptians, Hittites and Babylonians used bronze weapons, made by combining copper and tin. Tin was so scarce that only the richest and most powerful groups could access it and the bronze weapons it allowed. (Tin was so rare that it was being transported from Cornwall to the Middle East more than 1000 years BC.) The scarcity of tin gave these empires a near-monopoly on violence. Iron weapons are not actually any better than bronze ones – they're not stronger or easier to work with, as I'd assumed. They're just much, much easier to mass produce once you have bloomeries, which are the first furnaces that were hot enough to melt iron in a way that made them useable for weapons. Bloomeries meant that iron ore, which was widespread, could be used to make weapons anywhere. That led to warlords, bandits and city-states rising in a massive decentralization of military power, which the Bronze Age empires were unable to resist. After the invention of the bloomery, skeletons show a measurable increase in weapon-inflicted injuries. 'Destruction layers' with bones and signs of burning have been found from this time at archaeological sites including Troy. Ancient writers – including Herodotus, Ovid, and the writers of the Old Testament – believed the invention of iron had unleashed a new age of violence. New at Works in Progress, WEAPONS OF MASS DECENTRALIZATION: how iron itself brought about the Bronze Age Collapse. And what technologies could do the same today? worksinprogress.co/issue/wea…
1
396
Asked Opus 5.5 to make a web app from my gym trainer's exercise sheets. Started when I arrived, v0.1 was ready before I got warmed up, continued iterating over the session. (Figuring out data sharing was the only tricky part.) pasky.github.io/gym/#/view/p…
1
1
424
the new bipolar world
People keep not fixing this flow chart. Here you go.
5
363
Petr Baudis retweeted
I used to debug code now I type this and hit enter
179
1,086
23,500
516,157
...and there goes another idea in my backlog! They keep falling faster and faster :) Notably, the big unlock here was Astra also building its own harness, AND figuring out what exactly the best harness is, beyond a few high level ideas. (Subtweeting a few copes on my TL.)
In all seriousness, this is a startling achievement for GPT-6 Astra. kenforthewin.github.io/blog/… (This is GPT-6 Astra beating Nethack on its 3rd try. Nethack is the original roguelike and one of the most famously hard games of all time. I have played a lot, and I've never ascended)
2
7
852
Astra few days into an autoresearch loop vs. Opus 5.5 after ~6 hours of work, btw
> wife: "if you ever want to try and convince me that AGI ruling over humanity is around the corner by making a 3D model out of a picture, here is the picture" > me: "ez" > [week 3] Sol just gave up for good. Fable is still plowing on (currently failing on a reverse engineering ToS filter). (might make a good bench, idk?)
6
3
273
45,660
Petr Baudis retweeted
lmao what is this? now we know why they are so cheap
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others Pricing is approximately half that of GPT-5.6: Sol drops from $4/$20 to $2/$10 per million input/output tokens, and Luna from $0.20/$1.20 to $0.10/$0.50, with the same 90% discount for cache reads and 25% premium for cache writes. Key takeaways: ➤ Halves Cost per Task: GPT-6 Sol (max) costs $1.06 per task to run the Artificial Analysis Intelligence Index, ~50% less than GPT-5.6 Sol (max) at $1.99. GPT-6 Luna (max) costs $0.07 per task, ~60% less than GPT-5.6 Luna (max) at $0.18. This is driven by the price cut, as both models use slightly more output tokens per task (31k vs 29k for Sol, and 51k vs 41k for Luna). These two releases allow OpenAI to capture a significant portion of the cost efficiency Pareto frontier. ➤ In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task. ➤ Significant reduction in hallucination: Both models hallucinate less in AA-Omniscience, our knowledge and hallucination benchmark. GPT-6 Sol (max) cuts its hallucination rate from 92% to 60% and GPT-6 Luna (max) from 93% to 77%. Sol achieves this by declining to answer more often: it attempts 83% of questions vs 99% for GPT-5.6 Sol (max), which cuts wrong answers by about a quarter but also lowers accuracy 5 points from 59% to 54%. Luna's accuracy is broadly unchanged at 44% vs 43% while it answers fewer questions. On the AA-Omniscience Index, Sol improves from 22 to 27 and Luna from -10 to 1. ➤ Mix of improvement and regression across evals: Beyond AA-Omniscience, both models improve in AutomationBench-AA (Sol 62% vs 60%, Luna 53% vs 50%) and Terminal-Bench 4.0 (Sol 44% vs 40%, Luna 13% vs 12%). However, we observe regressions in two key knowledge work evaluations. In GDPval-AA v2.1, our benchmark adapted from OpenAI's dataset of economically valuable tasks across 44 occupations, Sol drops ~100 Elo points and Luna ~75. Luna also drops ~45 Elo points in AA-Briefcase v1.1, while Sol is level. AA-Briefcase v1.1 is a private evaluation across multi-week knowledge work projects, with thousands of input files. Our team has manually inspected hundreds of model outputs: the regressions tend to be driven by reduced presentation quality and deliverables that omit rubric elements. Congratulations @OpenAI and @sama on the launch!
100
30
1,584
323,160
Astra is washed btw
that is pretty good
1
1
865
guys guys Sonnet and HAIKU not dead??
Replying to @mikeyk
We’re also launching Claude Sonnet 5.5 and Haiku 5.5 in the coming weeks. These models will come with many of the same improvements to performance, efficiency, and safety. With that, happy building everyone.
1
444
Petr Baudis retweeted
Is Xiaomi's the most vertically integrated AI lab in the world? Most of the labs are predominantly software houses. And the sheer variety of consumer products probably beats X, which would be the obvious runner up.
From a research question to candidate materials. Xiaomi MiMo-V2.6-Pro worked with Xiaomi’s materials researchers to explore new MOF designs for capturing PFAS — reviewing literature and patents, proposing and checking design hypotheses, and running computational “dry-lab” experiments to identify promising candidates for further validation. The team estimates a 10× productivity gain, shortening the R&D cycle from one month to 2–3 days. As Prof. Jinhu Dou of Peking University put it: “In my view, its work on this project—from literature review to materials design and computational evaluation—was on par with that of a well-trained doctoral researcher.” Read the full case study: mimo.xiaomi.com/blog/mimo-v2… #XiaomiMiMo #XiaomiMiMoV26 #AIforScience #XiaomiAI
1
1
1
653
Petr Baudis retweeted
even anthropic is giving a banked reset. nothing stands in front of the power of the free markets.
Replying to @claudeai
One more thing: we’re increasing five-hour usage limits on Pro, Max, and Team plans. We’re also providing subscription users a rate limit reset, which you can save and use whenever you choose.
6
2
70
2,412
Petr Baudis retweeted
bro just dropped Gemini's system prompt
you are going to fail, so fail while daring greatly
123
462
12,853
573,278