co-founder, board @coreauto - and continuing to aspire to understand deep learning.

Pinned Tweet
It turns out multi step backpropaganda is better. paper has a beautiful way of improving backpropagation. One iteration cleanly gets us backprop, multiple iterations get us a preconditioned update.
5
12
231
169,774
Yesterday was a very good day of new technical milestones at @coreauto You can feel the gradients wants to learn.
4
77
3,973
Bahdanau is 🐐 He cooked edge transformers years before we got to 2-simplicial as well.
yeah this checks out my first few months at periodic, the weekly midtraining meeting was: 1. guy who trained the first trillion param LLM 2. guy who invented the attention mechanism 3. me you learn very quickly surrounded by this density of talent. we're hiring btw.
6
5
82
13,427
Legend, very early into things
9
1,188
Sometimes I don’t understand neoclouds sometimes I am grateful they exist.
2
13
2,460
Every time life throws me lemons, I ask my agents to make the best deeply stacked lemondade.
2
16
2,134
Ranthambore Tigers 🐅 Where every tiger walks like it owns its kingdom. Wild. Majestic. Unpredictable. Ranthambore
4
10
124
2,230
I am definitely pulling the average down @coreauto
I asked people which neolabs under $10B have the highest talent density: 32 votes: Core Automation (@MillionInt, @_arohan_) 26 votes: Periodic Labs (@LiamFedus, @ekindogus) 20 votes: Flapping Airplanes (@spectorb, @amspector100, @aidanmantine) 18 votes: Standard Intelligence (@G413N, @devanshpandey) 8 votes: Ricursive (@annadgoldie, @Azaliamirh) 5 votes each: - Ineffable (David Silver) - Mirendil (@bneyshabur, @HarshMeh1a, @shayan_, @tararezaeikh) - Recursive (@RichardSocher, @_rockt, @timshi_ai, @josh_tobin_, @CaimingXiong, @jeffclune, @tydsh, Alexey Dosovitskiy) 3 votes each: - Applied Compute (@ypatil125, @rhythmrg, @lindensli) - Isara (@ezhang7423, @hegasz) - Prime Intellect (@vincentweisser, @johannes_hage) 2 votes each: - Elorian (@AndrewDai, @yinfeiy, @SethInternet) - Goodfire (@eric_ho, @DanJBalsam, @banburismus_) - Inherent (@tantumscollins, @edwardfhughes, @LouisKirschAI, @kallyaleksiev) - Trajectory (@rronak_, @QuantumArjun, @MichaelElabd) - World Labs (@drfeifei, @jcjohnss, @BenMildenhall, @chlassner) People couldn't pick a company they work at or founded. Disclosure: I'm a small investor in Applied Compute, Factory, Standard Intelligence, Trajectory and Wafer. I didn't vote. Continued:
4
71
11,603
Building a company, particularly as they call it “neolab” is hard work, innovation is only way forward, and it is the best way as well. I think it’s all about the drive and motivation and ability to work in compounding ways, more than talent alone.
1
80
4,728
On that note, my cat has been waking me up every 5 am by meowing into my face, is there any simple way to fix this?
2
9
1,298
JFC, I am not operating the coreauto account. * just for clarity
Replying to @_arohan_
rohan poasting on the coreauto account then replying using his own account XD
1
28
4,669
rohan anil retweeted
Do you really need to finish every rollout in RL for LLM? Not if you make your critic better—and trust it! We propose Actor-Critic with Action Chunking (AC2): use a learned critic to score chunks of tokens => only need partial rollouts ⇒ train faster than GRPO!
21
90
631
70,636
The biggest difference between RL and PT.
Is it 📈📈📈 or 📉📉📉 ?
1
80
8,944
I tested out automated memegen at CoreAuto mimicking Googles famous memegen where I spent a lot of time since 2013. It created really funny things at-least according to me, and even used many agents, reasoning tokens to make it funny. Some people like it as a newsfeed but generally thought it sucked joy knowing it was generated. I tried make it more curated, but wasn’t as helpful
In this era of artificial intelligence, it is becoming urgent to distinguish human art from what machines produce. There is an ontological difference, even before an aesthetic one, between art and what a machine can generate through statistical calculation based on millions of images created by others. Algorithms lack the spark of humanity. For this reason, the Church wishes to renew an alliance with artists and cultural institutions to safeguard our humanity.
3
26
5,955
Things like these
3
16
2,686
I wonder why I find the idea that models having personhood illogical but other smart people think it is. I am curious to see the world with their eyes to see why I am wired differently.
18
52
6,431
Really nice work!
Kev 1.0: an open weight family of decision models you can train and run on your own. • New Kev-27B + updated Kev-9B • 64k-token document window on Kev-27B • TypeSafe SDK compatible • New fine-tuning + deployment skills so you can train Kev on your own data
28
6,902
The more interesting thing I have heard this week is that folks at some labs are mining alpha from my posts, and having their agents work on it. It seems like a great idea, so I am starting to do that as well.
9
5
187
15,626
Residual streams is the wrong perspective. It is some form of memory but we should have called it residual tape. But hold on, it cannot be a residual either its additive, any component can be larger than the other tape library sounds like the right term? Anything else?
9
3
91
8,975
At work when we have conflicts, we take our modded RC cars and use that to race each other, and see who wins. This is how compute is allocated.
7
80
5,027