monitoring frontier ai and cultural pulse in real time check trackingsingularity.com looking for work around rl envs and qa, prev ai eng and consulting

bangalore, india
Based in India
Pinned Tweet
new blog on my explorations @ prime residency "Honey, I Looked at the Data of a Frontier Benchmark and Found some Issues: Porting Agents' Last Exam Linux CLI subset to Verifiers v1" is up now. got to work with florian (@xeophon) as my verifier on this. sankalp.bearblog.dev/porting…
8
20
172
36,942
last 24 hours on trackingsingularity[dot]com what's the deal with sweden? some vpn shit?
1
1
13
949
about to eat a dosa in surat, wish me luck
7
37
1,348
Is there cheese filling?
1
3
124
luckily it didnt have
3
51
What’s one thing that’s missing in codex that you wish we had?
8,163
123
5,716
1,500,670
gpt 6.1 astra
37
955
crazy
you’re invited to the super intelligence terminally online luncheon. which seat are you choosing?
3
32
2,976
Is roon sam alt?
1
64
people may not take you seriously but you can take yourself seriously anon. its required to be playful.
3
1
24
879
you must figure out how to curiosity maxx. good questions and claude go a long way.
1
2
42
1,171
never leaving this bird app, i fucking love capitalism
81
2,523
at thailavaa, ok its good
234
the chai must go on
2
21
893
bhau aaj akele kyu?
1
1
113
not in blr rn
1
28
join vals! PS: we also have a fellowship program
Prediction: universities will get priced out of AI evals work We just spent $50k yesterday benchmarking 40 long horizon coding tasks As we design harder tasks for agents, trajectories are getting longer and need more tokens. That means substantially more inference and compute. How is this economically feasible for academics working off university grants?
2
36
6,557
do u folks give remote contracts. will prolly dm. (please checkout my pinned blog for some recent work)
222
chai must go on
1
11
357
somebody needs to write a short story about mind-uploading, where a guy spends his whole life developing the technology, and right at the very last moment, as he's about to upload himself, he realizes he has just committed a very convoluted form of suicide
8
30
1,967
i sometimes wonder why nyc is a creative ppl hub while being an urban jungle. i havent been there yet so perception totally based on media and listening to friends. i think a blog can be written on this topic.
2
7
1,173
sankalp retweeted
One of the pivots I did over a year ago when I started as Faculty @LTIatCMU and @CMUEngineering was towards AI for Molecule Design and Drug Discovery, and today we are releasing a blog post + our @PrimeIntellect and @harborframework environments for our benchmark SMDD! We find that: 1) Harness optimization can go a super long way in scientific tasks, and we are still heavily bottlenecked by the model capabilities rather than the knowledge, 2) Automating the harness optimization can work on some cases and not others! Amazing work led by @SureshRaghu07, w/ @KevinH1119568 and @aviral_kumar2
New blog: Is human taste overrated in harness engineering? An agent harness is the system around a model. It decides which tools the model can call, what evidence it sees, and what it remembers between steps. Most harnesses are still designed by human taste. Run the agent, read where it failed, change the harness, try again. That loop is increasingly being automated, with a strong model rewriting the harness itself. So we asked how far that can go. We ran Qwen on three drug design tasks from SMDD-Bench and compared a manually redesigned harness against one produced by automated harness search with Claude in the loop. No single approach won everywhere. The manual harness pulled clearly ahead when the fix was changing what the agent sees and remembers. Automated search did better once that was in place and the remaining problem was how to search the chemistry itself. The hard part was rarely implementing the fix. It was diagnosing what kind of failure we were looking at, and that is where human taste still mattered. Work with @KevinH1119568 , @aviral_kumar2 and @niloofar_mire Read more 👇
7
21
197
22,890
pacing the frontier with @sachdh
4
1
36
1,464
dayum
2
196
why is thinking machines not in here. not considered neolab because they make lots of revenue? fwiw they have crazy talent density.
bro we literally laugh over 80% of the time
5
1
78
9,621
they did prime intellect so bad😭
1
11
1,279
Replying to @dejavucoder
“under $10B”
887
Replying to @dejavucoder
“under $10B”
1
5
1,224
ah i missed
443