Pinned Tweet
ok it's done and released on goatremote.com Mac & Linux both work now. the remote feels great latest versions of @OmarchyLinux, Ubuntu & Kubuntu are supported here's Omarchy TV running on a raspberry pi:
ahhh just made a breakthrough, linux version coming this will be the best smart TV you can imagine, with the best remote hardware that exists, now on fully open OS. fixing this civilizational issue once and for all
20
33
620
52,626
murat 🍥 retweeted
🎯 Learning is most effective at the frontier of capability: problems that are too easy or too hard teach nothing. For LLM reasoners trained with GRPO this is literal: problems the model always or never solves give zero gradient. We introduce Frontier Learning👇🧵
23
53
859
75,322
intuitively makes sense, by rewriting training data to optimize for surprise only where it matters the backprop is given much cleaner signal. with this line of thinking one could expect GRPO to not work that well when using N human provided samples instead of on policy generated
SFT is not dead! 🥳 We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and forgetting less than RL and OPSD. 🤯 Following our prior work on reasoning with sampling, we now introduce sampling to the posttraining stack. 1/n
4
35
3,198
not just for creatives.. the selfish channeling of impulses and selfish intrigue in own perception , makes people very interesting and charming
favorite reminder for creatives, by Björk
6
49
2,921
in nyc for a bit
4
726
murat 🍥 retweeted
solving mechinterp and making embedding search work reliably are the same task
3
1
15
1,918
murat 🍥 retweeted
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
84
459
3,549
736,719
i (i mean gpt) ported the original Soldat to web soldatweb.com requires login with x so we know who is who. sorry the bots are too good i should nerf them
6
839
murat 🍥 retweeted
Replying to @allTheYud
LLM's still have the completion engine soul.. in the rare occasion, a pivot token, (here likely "freed"), they double down on the wrong thing. increasingly rare with post training BUT ALWAYS THERE all you need for misalignment events is for one to start off a feedback loop (like the agents message boards) where it virally spreads to other contexts i really think the CORE of the problem hasn't changed since gpt-2 or whatever. just less common. when alignment is a "march of the .9's" statistical problem, sheer volume of use is pretty much ensured to create tail events
2
2
11
1,122
so Jev turns out to be quite good at jailbreak detection. it beats gpt luna for example. good cheap solution for pre-screening prompts, and this is just super dumb initial attempt. i can prob get this close to 100% for most known jailbreaking patterns
Replying to @allTheYud
am i the only one who thinks it's likely very easy to build a cheap small classifier that detects jailbreaks? they rarely look like real requests if ever
7
10
266
15,666
idk if it's the only but it's surely the most guaranteed and therefore scariest to me
the only misalignment AI xrisk that is happening right now is an insidious infection of our capital systems
2
1
4
1,591
ok this is genuinely a new thing instead of answering with text, answers questions in parallel by either giving True/False responses, or assigning probabilities to up to 255 choices you provide. likely a great tool for LLMs to use. i wanna do constraint propagation experiments
Replying to @CompleteSkeptic
The gains aren’t free: Jev can't generate text Comparing Jev vs LLMs side-by-side makes the trade-off clear Fun fact: replacing sequential computation with parallel is the same way Transformers leapfrogged RNNs
14
5
559
81,973
update nvm it's not a new thing; @rysana did this ages ago
1
5
332
first Jev test: pretty good but 90% agreement rate with verified gemini flash workflow after some tuning it would prob get where i need but not a drop in replacement atm
8
1
67
12,560
murat 🍥 retweeted
Pre-orders for our tactile glove are now open firstcontact.xyz
v1 of our tactile gloves!
6
7
30
3,659
pretty proud of humans for how we master a topic in 200 kWh
3
40
2,135
superintelligence with its trillion watt hours can kiss my ass
5
370
murat 🍥 retweeted
> be me > bottomless sandbox supervisor
1
2
14
901
im not worried about self replicating robots because the manufacturing world is tied enough to the meat space legal and funding system to be caught early not worried about nano self replicators consuming earth bc i don't think you can think your way into that kind of tech tree not worried about digital hacking bc it's ez to recover from complete loss of digital records. hacking of physical systems is very scary but also generally localized i'd say. hacking nukes idk extreme example i'm a bit worried about engineered viruses but i think we'll recover from them too i'm mostly just worried about extreme changes to culture. i don't think people realize how much of the economy is a game for us to achieve collective consensus and not a real need. we're introducing a new economic doppleganger to fuck with us at every step, pretending all the work needs doing for reasons other than finding social balance. we're really gonna struggle with consensus on allocating true scarcities.
5
4
37
2,045
even the safest AI's lead to a future i find deeply unsettling at this point after seeing how the first three years went
4
7
622
> be me > bottomless sandbox supervisor
1
2
14
901