co-founder/ceo @neosigma_ai. ex-early @p0 | phd dropout @mit

San Francisco, CA
crazy how @alexandr_wang went from hardcore b2b to hardcore b2c And he is killing it all!
11
763
The interest and enthusiasm for evals track at @modal runtime conference was insane!! The queue was literally till outside the conference. So great to see more companies measuring what matters. Thanks for such an amazing curated conference @modal @akshat_b
1
3
25
1,318
Out of all the use-cases for personal agents, crazy that he talked about this!
My conversation with Noah Shinn (@noahrshinn), founder of Instinct. Noah is building a personal AI assistant. It's still invite only, has spent nothing on marketing, and is growing roughly 10% A DAY. This is his first long conversation about the company. We discuss: - Why Instinct doesn't have an app - Buying compute months ahead of exponential demand - How users learn to trust it with a credit card - Safety and security - Agents coordinating with other people's agents - Instinct's business model - Apps built on consumer inertia - and more Enjoy! Timestamps: 0:00 Intro 4:11 What people are using AI agents for 15:07 Rethinking travel, reservations, and the internet 22:43 Trust, privacy, and personal data 27:50 The business model behind Instinct 38:04 How existing businesses will adapt 47:55 Designing a personal assistant people love 53:15 Growth, compute, and competing with Big Tech 1:11:44 What’s next for Instinct and personal AI
1
7
1,199
Apparently, its one of the history’s biggest deals being signed by women on both ends of it! great to see them coming together! role models @LisaSu @drfeifei
We are excited to announce that World Labs is joining @AMD. The research and technical breakthroughs we have achieved since our founding in 2024 have given us a clear vision for AI’s potential to solve problems in the spatial and physical world. Accelerating the future of spatial and physical intelligence requires scaling our efforts, scaling our reach, and getting closer to the hardware.
16
1,608
Such exciting news from @ListenLabs ! And we’re so proud to have been serving them as one of @neosigma_ai's one of earliest customers. I still remember walking over to their office with @ritvikkapila, and a laptop in hand, with a local demo and a vision: that production agents shouldn’t freeze after ship - they should continuously evaluate, learn, and optimize from real user traffic. Listen believed in that early. They became one of our first customers, and we’ve been fortunate to serve them ever since. From those early demos to every piece of feedback their team and Tobias has given us, @ListenLabs has shaped @neosigma_ai tremendously. They pushed us to make our product better everyday. Taking agent data from production, turning failures into high-fidelity evals, and using that signal to improve the harness so agents get better at the outcomes that matter to them! It’s rare to find a team that cares this deeply about quality while moving this fast. Working alongside @itsalfredw , @florian_jue , Tobias, and the @ListenLabs team has been one of the privileges of building NeoSigma. We’ve been fortunate to help power evals and improvements for their agents as they’ve served each one of their customers like @AnthropicAI , @Microsoft and other giants with their agents. Looking ahead, we’re excited to keep supporting Listen as they take that vision further - now with @salesforce and serve even bigger customers. To @ListenLabs. Congrats, Alfred, Florian, and the whole team. What you’ve built is special and it’s been an honor to build alongside you all!
Listen Labs is joining @Salesforce! A year and a half ago we launched Listen, an AI platform for understanding customers, and it’s been a crazy ride ever since. Today, we work with some of the largest companies in the world, including Microsoft, Anthropic, and Sweetgreen. What started as an AI interviewer is now a full platform for understanding customers. Listen finds the right people, interviews them, analyzes what they say, and even simulates how they'll behave. Listen’s growth quickly accelerated and while we were raising our next round, we met @Benioff. Marc is a hero of mine. I've taken countless ideas from Behind the Cloud and watched him framemog every AI CEO on the planet. We're honored to get framemogged next. With Salesforce, we can bring Listen to every company in the world, much faster. To our customers, our mission remains the same.I stay CEO, @Florian stays CTO and the same team will keep building the product you rely on, now inside Salesforce AI Labs. Thank you for believing in us early. To our team, families, and everyone who believed in us before there was much to believe in, thank you. These years have been the most fun years of my life, and it’s still only just the beginning for Listen. Now back to work. 🚀
5
3
49
5,488
Every agent trace is an opportunity to make the next interaction better. But when verification is expensive, teams inspect a small sample and miss the failures hiding in the long tail. At @neosigma_ai, we turn those failures into evals and feedback for improving agents. Our latest blog explores how Jev makes that feedback loop cheaper and faster to run at scale. Across 10,000 verifier decisions, we measured 6.2× lower cost and 34× faster median decisions with Jev compared with GPT-5 Mini, alongside 88% versus 85% agreement with human labels in a blinded 100-trace audit. read our full blog here: neosigma.ai/blog/verifying-a…
4
3
24
2,686
This! Evals are one of the gate (and the only trusted gate) for diffusion of AI in the enterprise. If you want to own your frontier labs grade evals from your prod traffic, come talk to us @neosigma_ai
You can’t automate what you can’t measure. This means that evals are one of the gates to diffusion of AI in the enterprise. We can test our deterministic processes through software, but most enterprises have no useful way of understanding how their non-deterministic processes are working today. Specifically the work that agents are doing for them. Evals are mission critical for enterprises adopting AI because you have no other way of knowing what’s working, what’s broken, what changed, what improved, what you can do more of, etc. if you don’t have a good sense of how agents work in your environment today. All changes, upgrades, and deployments are downstream from good evals. Not only are we going to get vastly more domain specific evals over time for the labs and across the industry, but every enterprise will also need a clear sense of how agents are performing in their environment as well. Huge opportunity.
2
12
1,737
Gauri Gupta retweeted
At the IIA Silicon Valley AI Summit, I spoke about agent evaluation, optimization, and long-horizon reasoning; how to understand where agents fail, learn from production feedback, and continuously improve their performance over time. This is core to our mission at @neosigma_ai: helping teams understand, evaluate, and improve agents from real-world experience. Thanks @imaginationxyz for bringing the conversation together!
2
18
539
This has to be the most unconventional speech of its time 🙌
Jev Founder, Diogo Almeida (ex-OpenAI): "The next era is not the Claude Code or Codex era, they are still part of the assistance era with human in the loop - JEV is what comes next for LLMs x200 faster, x400 cheaper, 0 hallucination, no human in the loop - that's JEV, this is how LLMs will look like" in 36-minute tech talk, Jev Founder explained why RLHF isn't a thing anymore and how modern LLMs will be built this talk is worth more than a Stanford Machine Learning degree watch today no matter what, then learn how to become a Jev Engineer in the article below
11
235
80,739
and build your benchmarks directly from YOUR prod distribution. happy to support that @neosigma_ai
if you are building a product using AI, you should be spending >25% of your time making benchmarks and trying to get the model labs to care about said benchmarks easiest path to accelerate your progress as a company
2
2
14
2,474
Gauri Gupta retweeted
if you are building a product using AI, you should be spending >25% of your time making benchmarks and trying to get the model labs to care about said benchmarks easiest path to accelerate your progress as a company
257
135
3,532
1,134,942
my feed still all jev!
4
1
13
826
Gauri Gupta retweeted
Our Co-Founder, @ritvikkapila, will be speaking today at @imaginationxyz , joining “From Reasoning to Action” panel at Google Bay View. If you’re attending, come talk to him to learn more about what we’re building at Neosigma. imaginationinaction.co/2609s…
3
7
279
evaluations of models and agents that are doing real work on behalf of humans is going to super critical for both understanding what they are capable of and how we can safely diffuse AI into every sector of work. companies building powerful agents will also soon need to do such independent evaluations to ensure safe and efficient use.
on the idea of evaluators: think it's important that we have a distributed ecosystem of indepedent evaluators. the more eyes and people with distributed skill sets the better. it would be a good idea to fund several efforts on this.
3
14
1,466
Gauri Gupta retweeted
Inference roadmap I gotta follow: 1) Watch this course fully : piped.video/playlist?list=PL… (Covers everything related to LLM Inference+ Training but on high level) 2) Gonna read this resource: jax-ml.github.io/scaling-boo… ( for deep dive on topics) 3) Read PMPP(CUDA & GPU specific chapters) 4) Blogs and codebase of vLLM , SGlang, llama.cpp 5) Inference Engg by philip book for brush up on the concepts 6) Also these notes which i found randomly on X for brushup: x.lingyaoai.com/gauri__gupta/status/20… Along with it : contributions to the above mentioned repo for understanding depth + talking with folks building in inference here
14
72
663
61,296
Gauri Gupta retweeted
This playlist is a great place to start before going deeper into LLM inference. It covers Attention, MoE & architectures, GPUs, CUDA & Triton, parallelism, and scaling at a high level. A very useful resource for building the fundamentals before diving deeper into LLM inference systems. piped.video/playlist?list=PL… Sharing it along with some notes I found very useful. Worth checking.
1
7
91
10,912
If you are excited about our mission @neosigma_ai, improving reliability in long-horizon agents and evals, @ritvikkapila will talking all about how we are doing it here. Come say hii! imaginationinaction.co/2609s…
Such a killer panel with @silasalberti, @CyrilGorlla, and @PaulBaier , going to be so much fun at @imaginationxyz talking about long chain reasoning!
2
12
1,922
Gauri Gupta retweeted
I’ll be speaking at Imagination in Action (@imaginationxyz) at Google Bay View on September 14–15, 2026, discussing the future of agentic infrastructure and how we can build more reliable AI systems. If you’re excited about our mission at @neosigma_ai, please reach out - I’d love to chat!
1
1
17
859
High signal Evals is sometimes all you need. - first defining what success looks like for your agent - building evals to then understand what your agents can do today, what is the frontier you are chasing - using these evals to also attribute the required amount of intelligence needed to solve the task (building model routing policies) saving costs - and hill climbing evals to optimize towards the next frontier of intelligence - evals are both test and train bed for the intelligence
evals ≫ harness optimization (fast learning) ≫ model post-training (slow learning) Most agentic tasks don't need slow learning loops at all. Even when the goal is post-training, you'd first need high SNR evals, then optimize the harness on the signal to get the performance up, and then "distill"* the resulting system behavior back into the weights. * Yes, I called it distill.
9
5
79
6,097