Rob Farlow retweeted
Who grades the graders? LLM-as-judge on short answers is well studied. Agent judges are not: they open the files, run the code, and investigate before ruling on another agent's work. Introducing Arbiter-Bench: 71 real agent runs that frontier agent judges get wrong.
13
16
52
3,505
Rob Farlow retweeted
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
116
193
1,935
635,406
Rob Farlow retweeted
America is banning AI in schools. China is using AI to create geniuses. Introducing Aristotle: The AI tutor that solves America’s broken education system. heyaristotle.com
414
462
3,351
1,415,586
cool graph on @mintlify of agent vs human traffic to docs pages over time. surprising it only passed 50% this june
1
12
677
oh wow didn't know grok bot uses opus 5 under the hood for computer use
4
1
29
4,890
this really shows the product ingenuity from the cursor/xai team. grok bot could have been made a while back it's just the innovation of having the permanently running cloud VMs and the grok bot interface
1
362
@jiaying_ai found harness info in /home/box/sand-host/host-main.cjs
1
2
737
Rob Farlow retweeted
If they count Sam Atlman's honorary degree from University of Waterloo, we no.1 But 7 is not bad for a small town in Ontario.
A first for a Canadian university - The University of Waterloo lands top 10 spot in global ranking for producing entrepreneurs kitchener.citynews.ca/2026/0…
3
2
89
11,886
I thought that software engineering was solved but after having to do some production debugging this weekend I realized its definitely not. You can still be a 10x engineer
1
16
1,461
I love the gamification of paraform showing me what percent engaged I am vs other hiring managers. feels like i'm climbing in a video game
9
1,257
prediction: nvidia will be a frontier lab by mid next year
1
1
14
844
grok 4.6 takes more than 4x as many steps as other frontier models on computer use tasks it is incredibly great at agentic work but sucks at clicking stuff on screens. almost every click it will have to try 4+ different times until it finally gets the right coordinates
6
27
2,661
*all models have 100% rate on these tasks and they are incredibly simple without any complicated reasoning
1
3
613
wait what how is x money paying 6% apy
10
1,205
i've recently been getting dozens of people reaching out to me to sell real enterprise data. seems like this market is exploding but I don't think people realize how much work is required to make this actually usable. this kind of data by itself is maybe worth 5% of the total value of an environment
2
1
33
3,162
anthropic has the best computer use model but the worst co-work product it always has trouble using my multiple browser profiles, I can only connect one google account at a time, and worst of all the google drive connector requires going byte by byte to upload stuff
3
18
2,212
grok bot is a simple improvement over openclaw (easier setup, don't need mac mini) and claude cowork (runs on a cloud vm) but it unlocks so much possibility It's now so easy to run customized agents to monitor things, do routine work, or to perform tasks which require waiting for 3rd party action such as a multi day email conversation
10
1,172
grok bot can install stuff in our office
3
2
21
1,772
it even messaged with him to update the time making sure it didn't conflict with my calendar
1
1
392
claude's analysis of 🍓🍓🍓's claims from the past 2 years
‼️huge ssi news. ilya is about to take his first tentative steps out of the age of research and back into the age of scale. it’s time to smell what ssi is cooking. ssi have built a small reasoning engine that can compete with much larger training runs because his data is better curated to meta learning. but, more importantly. we’re about to step into the era of TTT (test time training. gradient descent happening in real time to solve your problems). so instead of a context window you get actual learning. and because it’s so sample efficient it can be trained on hard to verify tasks that other paradigms can’t touch. everyone else’s weights are frozen, they struggle out of distribution. ssi have created something that has a bundle of knowledge but can truly learn in real time and use that to your advantage. current approaches are trying to hack their way to ‘learn’ with memory tricks, this thing will updates its weights, remember key lessons, and finally feel like a human level reasoner. this is a huge paradigm shift from the king. early results are very impressive. we can stop watching memento on repeat. it’s learning all the way down, the descent is real.
1
14
1,271