Sane + 🌶️ takes in an insane AI world... AI capabilities researcher: co-created RLHF/ChatGPT @ @openai now trying to right the wrong 🤭 (ceo @typesafeai)

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
4,048
8,316
76,591
39,907,009
late to reading this, but FUCK YES to playing with more shapes for intelligence! please play around with it more! append-only chat is not the final interface!
Wrote a short blog about the "shape" of language models, and the tradeoffs they may present in the future. I genuinely think it's a valuable research direction to start thinking about now, especially w.r.t. harness design. alexzhang13.github.io/blog/2…
20
25
424
27,283
my favorite plot here was on consistency! we worked super hard to put all sorts of properties we feel are necessary for intelligent automation into the models this is also just the beginning ! 🤘
11
9
171
14,817
Diogo Almeida retweeted
How to build a reliable risk agent without a frontier model ($0.02/sweep): • @youdotcom search • @typesafeai Jev for typed judgments • @QwenDevs for proposals & synthesis • MCP for integration Full architecture & cost breakdown below 👇
Article

Building a Reliable Background Agent Without a Frontier Model

Reviewed by @typesafeai I ran three live, escalated risk sweeps through this system. Each sweep made six calls to Qwen3.8 27B through OpenRouter: five tool-calling turns and one report-synthesis

22
14
109
45,613
not you too jev!
Jev is now in the AI SDK for Python. To test it, we ran two experiments: ▪︎ Detecting Python vs. English as text is typed ▪︎ Writing Python, one decision at a time 𝚞𝚟 𝚊𝚍𝚍 𝚊𝚒 vercel.com/blog/jev-for-pyth…
15
3
160
25,352
Diogo Almeida retweeted
Inspired by @EGafni’s Twitter thread on combining Jev with PageIndex. We show how to build long-document search with @typesafeai Jev + PageIndex. No vector database. No embeddings. Open source: github.com/VectifyAI/jev-doc… 🧵👇
10
23
169
94,344
me to our GTM guy: do enterprises want benchmarks him: no, the fast ones have already benchmarked their internal use cases. the slow ones are copying the fast ones. he's onboarded a quarter of the fortune 500 already 🥵
52
33
1,003
82,927
super excited for this - should make workflows much easier to build!
Replying to @typesafeai
@typesafeai's Jev model just landed in n8n. Jev’s job isn’t to write, it decides: returning a choice + confidence you can use for branching, sorting and routing in your workflows. Think of it as a smarter If/Switch for the fuzzy stuff. bit.ly/4hFjS8A
17
10
119
18,608
Diogo Almeida retweeted
Jev is a phenomenal reranker. At 75% of the previous reranking cost, it improved our main-app search by 30.64% in matches per 200 candidates on searches recruiters actually run. Most recruiting benchmarks measure precision and accuracy across just 10 candidates. That isn’t representative of real recruiting and sourcing workflows, where you may need to source hundreds of candidates to make a hire. We’ll release the full benchmark soon. We're excited to work with @hmartenjoyer, @mathfax , and the rest of the @typesafeai team to continue to bring Jev to more Wrangle features. Stay tuned 👀
5
6
37
21,404
Diogo Almeida retweeted
Is no one gonna talk about the fact that the surname in reverse reads "No Slop"
44
188
2,849
123,057
Diogo Almeida retweeted
We ran agents on a site with and without Jev to see whether a decision model makes a site more accessible and usable for agents. Across browser-use, WebMCP and NLWeb: 3.6× faster on average (up to 4.3×), 7.7× cheaper on average (up to 13×). Faster in every one of the paired runs. Task success held. 240 runs, full method and data: ora.ai/blog/evaluating-jev
7
8
37
21,152
she's roasting me for taking 2 years AND leaking stuff on podcasts 🤦
who tf is leaking our product roadmap to MY MOM
45
6
530
80,540
who tf is leaking our product roadmap to MY MOM
42
13
800
117,159
begun, the clone war has jk, I love openai and think more competition and validation is great for developers! (assuming the model is good - plz make it good!) hopefully this is a sign for the future that building in a system one compatible way is the future
59
31
713
39,997
As a bonus, we're also releasing a 2nd dataset that we were planning to launch with before inventing the concept of workflow evals. Simply by evaluating on the 2nd one, we believe it to no longer be a valid indicator of true out-of-distribution generalization. (2/3)
2
1
70
10,756
p.s. prepare for more sick stuff next week!
3
59
6,654
It's been almost 2 weeks since we've launched! And due to overwhelmingly popular request, we're making the datasets released with evals.typesafe.ai easier to work with! We do this in the spirit of openness, but I am still anti-public benchmarks. In that spirit, we deprecate all datasets we evaluate on (internally or publicly). (1/3)
29
61
988
79,467
Diogo Almeida retweeted
We gave Jev 2,029 real phone calls. No transcripts or audio; it never heard a word. Our AI receptionist's calls were reduced to pure structure, meaning turns, tool calls, workflow stages and timing. During the calls, Jev made 38,012 turn-level forecasts at 118 ms median latency, reviewed every call with five typed questions and produced 10,145 answers in 26 seconds with 256 requests in flight. The experiment was zero-shot, with no fine-tuning or examples from our data. We compared Jev's forecasts with what actually happened in the EHR. By the halfway point, Jev could meaningfully separate calls that would book from those that wouldn't (AUC 0.78), and near the end it ranked them correctly 94% of the time. Even though Jev over-focused on visible errors our agent usually overcomes, it's still pretty incredible that it analyzed thousands of real calls in seconds for only $3.
104
137
2,144
367,912
Diogo Almeida retweeted
Three thoughts after playing with Jev: - this is way more than a simplistic classifier: you can pack a lot of complex state and do more than you’d think - it’ll take a minute to think about your problems in a non-LLM shape - not quite beating frontier LLMs in correctness most of the time so can’t quite switch yet but it’ll clearly get there - this is definitely a new model type we will all use by next year; I wouldn’t ignore it
40
24
545
51,674