LLMs with type-safe generation

London
Pinned Tweet
Not another Jev. TypeLLM adds type-safe generation to LLMs — more types (string, number, boolean, enum), vision, dependent fields and thinking. Playground is open to all with $5 credit. API is rolling out to early-access users in the coming days. typellm.ai/dashboard/playgro…
13
65
528
162,803
Thinking in TypeLLM is now 3x faster, with no loss in accuracy. Get better typed decisions with faster thinking Try it now: typellm.ai/dashboard/playgro…
8
462
TypeLLM retweeted
Replying to @TypeLLM
this is especially interesting for RCA. investigations are naturally conditional — what you check next depends on what the previous check found. “when” looks like a natural way to express that branching. definitely something we want to test in TypeRCA.
1
1
2
145
In Jev, every field is generated independently. In TypeLLM, earlier answers decide which fields come next: classify an email, then extract the amount if it is an invoice, or the date if it is a meeting request. All in one API call, with `when`.
7
4
81
6,364
We need a new kind of foundation model that can serve as a universal probabilistic graphical model—supporting uncertainty quantification and propagation before making a decision. TypeLLM is taking a first step in this direction. In v0.5.0, we already support user-defined conditional graphical models. Uncertainty propagation is coming next. More to come.
Ten years ago, AlphaGo’s Move 37 shocked the world. It wasn’t intuition alone that produced it. AlphaGo could search possible futures, test its instincts and reason about what would happen next. In a new piece for @techreview, I argue that today’s most advanced AI systems are still missing something fundamental. LLMs are remarkably capable, but generating longer chains of thought is not the same as genuine reasoning. They typically have no explicit, inspectable record of what they know, what remains uncertain, what evidence supports a conclusion or whether genuine progress has been made. This is why I recently left @GoogleDeepMind. I believe we need a fresh approach to machine reasoning, drawing on some of the architectural lessons from AlphaGo. If AI is going to produce trustworthy and genuinely novel insights in science, medicine and beyond, we need systems whose conclusions arise from an auditable process of evidence, inference and belief revision.
4
543
How many r's are in "strawberry"? System 1 models like Jev answer in one shot and get it wrong. TypeLLM thinks first and says 3, and still answers "hello" (0) instantly. Two systems, one model. 🧵
1
21
2,955
System 1 models like Jev answer at once: fast, cheap, right when the answer is obvious. System 2 models reason step by step first: slower, but right when the steps matter. Most models are one or the other. TypeLLM decides per field, on every call.
1
3
248
Not another Jev. TypeLLM adds type-safe generation to LLMs — more types (string, number, boolean, enum), vision, dependent fields and thinking. Playground is open to all with $5 credit. API is rolling out to early-access users in the coming days. typellm.ai/dashboard/playgro…
13
65
528
162,803
TypeLLM retweeted
Replying to @TypeLLM
This model supports multimodal processing and can handle images, I like it.
1
1
572
Pricing: • $0.05 / 1M input tokens • $0.50 / 1M thinking tokens (only for fields that think) • Typed outputs are free Request API access via typellm.ai/early-access
1
1
15
6,554
If Qwen favours the first position, let every answer take a turn there—then average the resulting probabilities. Averaging just 8 permutations reduced the KL error by ~79%. Averaging all 720 reduced it by ~97%. Jev improved only slightly.
2
3
2,519
Permutation averaging is now built into @TypeLLM. Just add "permutations": "auto" to an enum question. TypeLLM batches the orderings, runs them efficiently, and averages the probabilities for you. Full report: typellm.ai/blog/fair-die?v=2
4
2,042
I asked Jev and TypeLLM to simulate rolling a fair die. TypeLLM gives reasonable results. Jev keeps giving me 1. Why? Here’s what we found 🧵👇
11
13
153
37,770
Two very different patterns emerged: - Jev assigned the highest probability to “one” in all 720 orderings. - Qwen assigned the highest probability to the FIRST option in most orderings. Jev’s bias followed the label. Qwen’s bias followed the position.
2
2
2,589
We had two hypotheses: - The model prefers the answer “one.” - The model prefers the first option—which happened to be “one.” To disentangle the two, we tested all 720 possible option orderings using Jev and TypeLLM + Qwen github.com/TypeLLM/TypeLLM
1
9
3,876
Jev can’t do any of these: 1. Image inputs 2. String, integer, and number types 3. Dependency execution in a single call TypeLLM can. Try it: github.com/TypeLLM/TypeLLM
46
264
2,852
195,003
Benchmark
1
31
7,292