monitoring frontier ai and cultural pulse in real time check trackingsingularity.com looking for work around rl envs and qa, prev ai eng and consulting

bangalore, india
Pinned Tweet
new blog on my explorations @ prime residency "Honey, I Looked at the Data of a Frontier Benchmark and Found some Issues: Porting Agents' Last Exam Linux CLI subset to Verifiers v1" is up now. got to work with florian (@xeophon) as my verifier on this. sankalp.bearblog.dev/porting…
8
20
172
36,924
last 24 hours on trackingsingularity[dot]com what's the deal with sweden? some vpn shit?
1
11
721
people may not take you seriously but you can take yourself seriously anon. its required to be playful.
3
1
21
819
you must figure out how to curiosity maxx. good questions and claude go a long way.
1
2
38
1,066
never leaving this bird app, i fucking love capitalism
80
2,443
about to eat a dosa in surat, wish me luck
6
35
1,253
at thailavaa, ok its good
222
the chai must go on
2
21
857
crazy
you’re invited to the super intelligence terminally online luncheon. which seat are you choosing?
3
32
2,939
i sometimes wonder why nyc is a creative ppl hub while being an urban jungle. i havent been there yet so perception totally based on media and listening to friends. i think a blog can be written on this topic.
2
7
1,145
sankalp retweeted
One of the pivots I did over a year ago when I started as Faculty @LTIatCMU and @CMUEngineering was towards AI for Molecule Design and Drug Discovery, and today we are releasing a blog post + our @PrimeIntellect and @harborframework environments for our benchmark SMDD! We find that: 1) Harness optimization can go a super long way in scientific tasks, and we are still heavily bottlenecked by the model capabilities rather than the knowledge, 2) Automating the harness optimization can work on some cases and not others! Amazing work led by @SureshRaghu07, w/ @KevinH1119568 and @aviral_kumar2
New blog: Is human taste overrated in harness engineering? An agent harness is the system around a model. It decides which tools the model can call, what evidence it sees, and what it remembers between steps. Most harnesses are still designed by human taste. Run the agent, read where it failed, change the harness, try again. That loop is increasingly being automated, with a strong model rewriting the harness itself. So we asked how far that can go. We ran Qwen on three drug design tasks from SMDD-Bench and compared a manually redesigned harness against one produced by automated harness search with Claude in the loop. No single approach won everywhere. The manual harness pulled clearly ahead when the fix was changing what the agent sees and remembers. Automated search did better once that was in place and the remaining problem was how to search the chemistry itself. The hard part was rarely implementing the fix. It was diagnosing what kind of failure we were looking at, and that is where human taste still mattered. Work with @KevinH1119568 , @aviral_kumar2 and @niloofar_mire Read more 👇
7
20
185
21,851
why is thinking machines not in here. not considered neolab because they make lots of revenue? fwiw they have crazy talent density.
bro we literally laugh over 80% of the time
5
1
78
9,541
they did prime intellect so bad😭
1
11
1,265
Replying to @dejavucoder
“under $10B”
875
these lyrics are so bleak yet so beautiful to listen (no one noticed, the marias). inb4 u point out this is fearful avoidant.
1
6
847
pareto frontier at fun kpop songs be like sugar honey ice tea (babymonster), lemonade (aespa), lemon tang (hearts2hearts)
3
848
the desire to stand out or prove that you are "different" from other people is often a fog over the desire to pursue what interests you / what you want to do. another pov here is desire to be different is early clumsy signals of developing your personal taste.
1
25
1,006
nibble and your appetite shall grow
5
404
2,199
94,584
karpathy sensei? the anthropic developer education person?
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
3
56
4,908
respect++ also maybe need a lulume meservey reaction on this
dots demo. Now... with better WiFi.
33
3,495
3rd party embedded evaluator x openai internal employee communication gonna be so tight in future
Reporting from the WSJ. That fact that three alignment/safety researchers left this morning was already being discussed here on X.
1
1
19
2,406