Let your agents learn from experience in production.

San Francisco, CA
The next era of AI engineering is self-improving agentic systems! Really excited to share what we are building at NeoSigma! Self-maintaining agent systems represent a shift in how we build and operate software. We, at NeoSigma are building the infrastructure to support this feedback loop in real-world systems, helping teams capture failures, convert them into structured evaluation signals, and use them to drive continuous improvements in agent behavior.
We @neosigmaai @RitvikKapila are building the future of self-improving AI systems! By closing the feedback loop between production data and system improvements, we help teams capture failures, convert them into structured evaluation signals, and use them to drive continuous improvements in agent behavior. We show how our system works on Tau3 bench across retail, telecom, and airline domains. Agent performance on the validation set (with a fixed underlying model, GPT5.4) improves from 0.56 → 0.78 (~40% jump in accuracy).
1
2
23
12,645
NeoSigma retweeted
Such exciting news from @ListenLabs ! And we’re so proud to have been serving them as one of @neosigma_ai's one of earliest customers. I still remember walking over to their office with @ritvikkapila, and a laptop in hand, with a local demo and a vision: that production agents shouldn’t freeze after ship - they should continuously evaluate, learn, and optimize from real user traffic. Listen believed in that early. They became one of our first customers, and we’ve been fortunate to serve them ever since. From those early demos to every piece of feedback their team and Tobias has given us, @ListenLabs has shaped @neosigma_ai tremendously. They pushed us to make our product better everyday. Taking agent data from production, turning failures into high-fidelity evals, and using that signal to improve the harness so agents get better at the outcomes that matter to them! It’s rare to find a team that cares this deeply about quality while moving this fast. Working alongside @itsalfredw , @florian_jue , Tobias, and the @ListenLabs team has been one of the privileges of building NeoSigma. We’ve been fortunate to help power evals and improvements for their agents as they’ve served each one of their customers like @AnthropicAI , @Microsoft and other giants with their agents. Looking ahead, we’re excited to keep supporting Listen as they take that vision further - now with @salesforce and serve even bigger customers. To @ListenLabs. Congrats, Alfred, Florian, and the whole team. What you’ve built is special and it’s been an honor to build alongside you all!
Listen Labs is joining @Salesforce! A year and a half ago we launched Listen, an AI platform for understanding customers, and it’s been a crazy ride ever since. Today, we work with some of the largest companies in the world, including Microsoft, Anthropic, and Sweetgreen. What started as an AI interviewer is now a full platform for understanding customers. Listen finds the right people, interviews them, analyzes what they say, and even simulates how they'll behave. Listen’s growth quickly accelerated and while we were raising our next round, we met @Benioff. Marc is a hero of mine. I've taken countless ideas from Behind the Cloud and watched him framemog every AI CEO on the planet. We're honored to get framemogged next. With Salesforce, we can bring Listen to every company in the world, much faster. To our customers, our mission remains the same.I stay CEO, @Florian stays CTO and the same team will keep building the product you rely on, now inside Salesforce AI Labs. Thank you for believing in us early. To our team, families, and everyone who believed in us before there was much to believe in, thank you. These years have been the most fun years of my life, and it’s still only just the beginning for Listen. Now back to work. 🚀
5
3
49
5,482
Congrats to the @ListenLabs team! We’ve really enjoyed working with you and are excited for what’s ahead.
Such exciting news from @ListenLabs ! And we’re so proud to have been serving them as one of @neosigma_ai's one of earliest customers. I still remember walking over to their office with @ritvikkapila, and a laptop in hand, with a local demo and a vision: that production agents shouldn’t freeze after ship - they should continuously evaluate, learn, and optimize from real user traffic. Listen believed in that early. They became one of our first customers, and we’ve been fortunate to serve them ever since. From those early demos to every piece of feedback their team and Tobias has given us, @ListenLabs has shaped @neosigma_ai tremendously. They pushed us to make our product better everyday. Taking agent data from production, turning failures into high-fidelity evals, and using that signal to improve the harness so agents get better at the outcomes that matter to them! It’s rare to find a team that cares this deeply about quality while moving this fast. Working alongside @itsalfredw , @florian_jue , Tobias, and the @ListenLabs team has been one of the privileges of building NeoSigma. We’ve been fortunate to help power evals and improvements for their agents as they’ve served each one of their customers like @AnthropicAI , @Microsoft and other giants with their agents. Looking ahead, we’re excited to keep supporting Listen as they take that vision further - now with @salesforce and serve even bigger customers. To @ListenLabs. Congrats, Alfred, Florian, and the whole team. What you’ve built is special and it’s been an honor to build alongside you all!
8
429
NeoSigma retweeted
Every agent trace is an opportunity to make the next interaction better. But when verification is expensive, teams inspect a small sample and miss the failures hiding in the long tail. At @neosigma_ai, we turn those failures into evals and feedback for improving agents. Our latest blog explores how Jev makes that feedback loop cheaper and faster to run at scale. Across 10,000 verifier decisions, we measured 6.2× lower cost and 34× faster median decisions with Jev compared with GPT-5 Mini, alongside 88% versus 85% agreement with human labels in a blinded 100-trace audit. read our full blog here: neosigma.ai/blog/verifying-a…
4
3
24
2,684
NeoSigma retweeted
The cost of verification sets a ceiling on how much agents can learn from production. When you can only afford to evaluate a small sample, most of that experience goes unused. Cheap, fast verification raises that ceiling: more failures discovered, more representative evals, and more feedback to improve the model and harness. At @neosigma_ai, our work with Jev makes us optimistic about a future where learning from production is a continuous part of how agents operate. Across 10,000 verifier decisions, Jev was 6.2× cheaper and 34× faster at median latency than GPT-5 Mini, with 88% versus 85% agreement with human labels in a blinded 100-trace audit. full blog here - neosigma.ai/blog/verifying-a…
3
1
11
717
Production agents create more traces than teams can inspect. Those traces are the clearest record of where agents succeed and fail. That creates an evaluation bottleneck: semantic judging works, but costs too much to run broadly. We tested whether Jev, a model built for typed decisions, could reliably expand evaluation coverage. Full blog: neosigma.ai/blog/verifying-a…
2
1
10
322
Cheaper and faster only matter if judgments hold up. Human agreement was similar: 88/100 for Jev, 85/100 for GPT-5 Mini. The gap was inconclusive. The clearer split was failures: 25/29 vs 17/29. On repeated runs, verdicts changed on 3 Jev traces and 19 GPT-5 Mini traces.
1
1
31
Define success, evaluate production behavior, and see where agents fall short. NeoSigma helps teams evaluate agents at scale, discover long-tail failures such as partial completions, quiet policy violations, and plausible-looking failures, turn them into representative evals, and continuously improve agent behavior. Request a demo: neosigma.ai/waitlist
1
19
NeoSigma retweeted
and build your benchmarks directly from YOUR prod distribution. happy to support that @neosigma_ai
if you are building a product using AI, you should be spending >25% of your time making benchmarks and trying to get the model labs to care about said benchmarks easiest path to accelerate your progress as a company
2
2
14
2,474
Our Co-Founder, @ritvikkapila, will be speaking today at @imaginationxyz , joining “From Reasoning to Action” panel at Google Bay View. If you’re attending, come talk to him to learn more about what we’re building at Neosigma. imaginationinaction.co/2609s…
3
7
279
Our co-founder, @ritvikkapila, will be speaking at Imagination and Action (@imaginationxyz) at Google Bay View on September 14-15 about our mission at @neosigma_ai Please feel free to reach out if you're excited about our mission and we would love to chat!
I’ll be speaking at Imagination in Action (@imaginationxyz) at Google Bay View on September 14–15, 2026, discussing the future of agentic infrastructure and how we can build more reliable AI systems. If you’re excited about our mission at @neosigma_ai, please reach out - I’d love to chat!
6
448
We’re looking for someone to join our team and work across product, GTM, strategy, and partnerships. If you’re excited about building with us, we’d love to chat.
We're hiring a Founding Business Generalist to work closely with us across product, business, gtm, strategy, and partnerships at @neosigma_ai. We're looking for someone who understands technology deeply but is most energized operating at the intersection of business, product, and customers. You'll help shape the company from the ground up, tackling high-impact, ambiguous challenges across a broad scope in a fast moving environment. We're especially excited about former founders, people with VC experience, and high-agency early-stage operators. If you're excited about our mission, we'd love to hear from you!
1
7
908
We’re building infrastructure that helps AI agents learn from experience in production. We’re hiring across engineering, research, product, business, and more. If you’re excited about our mission and want to work at the frontier on these problems, we’d love to chat.
some news: - After 8 years running AI research at NVIDIA, @FidlerSanja teamed up with @ZGojcic and @HuanLing6 to build Veeda AI - a startup that’s scaling physical AI through interactive learning in simulated reality. The team just raised over $90M in seed funding - and they’re now hiring technical staff members across the board. - Cambridge computational neuroscientist @achterbrain just raised a $100M seed for Callosum, on a mission to make AI workloads cheaper. (This is just 6 months after the team came out of stealth with a $10.25M pre-seed.) Callosum is growing the team in London - including a Chief of Staff to the CTO, and a Head of Physical Infrastructure. - @gauri__gupta recently dropped out of her MIT Media Lab PhD and co-founded NeoSigma - infrastructure that lets agents learn from experience in production. She’s now hiring a Founding Business Generalist (SF) to work with her across product, GTM, strategy, and partnerships. - @TheRobertDoh (prev Citadel) and @GuptaPranit (prev Palantir) are building Motherboard Labs - agentic field operations for robot and autonomous machinery fleets. They're building out the founding team and in a16z speedrun right now! More here: a16zjobs.substack.com/p/open…
3
1
6
788