The context layer making AI 5x more accurate

Context Conference, Oct 28 👉
Packed house, a ton of learnings on Multiplayer AI from builders and stomach full of pizza. Thank you everyone for showing up and thank you for such depth. Insanely good quality builders, insanely good quality audience. @hydra_db & @AtlanHQ for the incredible partnership (+ demos)! @UseCorgi for being such an insanely wonderful partner for the space. What a star you are @jean_ette_li ! Can't wait to be back! Learnings coming up đź§µ
2
3
24
1,094
Absolute banger of an event - Multiplayer AI. Thanks so much to all those who presented, put together, genuinely the best, most fun, deepest, most tech savvy conversations in SF till now. @prukalpa @AtlanHQ @hydra_db Thanks for hosting @jean_ette_li and can I thank and tag Trudy the Corgi for vibes. So many endless applications of multiplayer AI 🤯 Would ChatGPT win it ? With their retail scale ? Could muse crack it ? What about Gemini ? Or could it be something else something non big tech. 🤷🏻‍♀️ So many questions, no correct answer, only contemplations and building and fun ✌️
4
4
20
1,657
Multiplayer AI runs on a spectrum, from your own team to the other side of a negotiation. Each point needs different levels of trust, scope and enforcement, and almost none of it exists yet. Here are the six big questions I keep hearing
Article

What needs to be true for Multiplayer AI to be real?

I've spent the past couple of weeks going deep with people: asking questions and trying to understand what needs to be true for multiplayer AI to be real. Two weeks ago, I wrote briefly that

4
4
17
3,745
Atlan retweeted
Banger questions on the day Dot released. Who are the coolest people on X working on multiplayer AI problems?
Multiplayer AI runs on a spectrum, from your own team to the other side of a negotiation. Each point needs different levels of trust, scope and enforcement, and almost none of it exists yet. Here are the six big questions I keep hearing
Article

What needs to be true for Multiplayer AI to be real?

I've spent the past couple of weeks going deep with people: asking questions and trying to understand what needs to be true for multiplayer AI to be real. Two weeks ago, I wrote briefly that

4
5
24
3,362
Atlan retweeted
Okay I just saw the demo lineup for this and it’s shaping up to be pretty cool. Excited to do this at @UseCorgi IRL in SF :) know anyone who should be there? Tag them?
The scene is set folks! Bringing together folks to make sense of Multiplayer AI, IRL in San Francisco on the October 1st :) Excited to announce and celebrate the partners making this happen @AtlanHQ , @hydra_db and @UseCorgi. Bring your questions, curiosities, demos, learn and share notes.
1
1
6
1,435
Atlan retweeted
I’m attending Context Conference by Atlan, the conference for teams teaching AI their business. The community is meeting to cover why organizations invest in context, real context architectures, open context formats, and a lot more. Join live on Oct 28 → atlan.com/context-conference…
1
2
52
Atlan retweeted
I’m attending Context Conference by Atlan, the conference for teams teaching AI their business. The community is meeting to cover why organizations invest in context, real context architectures, open context formats, and a lot more. Join live on Oct 28 → atlan.com/context-conference…
1
3
112
Atlan retweeted
I’m attending Context Conference by Atlan, the conference for teams teaching AI their business. The community is meeting to cover why organizations invest in context, real context architectures, open context formats, and a lot more. Join live on Oct 28 → atlan.com/context-conference…
1
2
132
Live from #FabCon in Barcelona: We're bringing governed context to Microsoft Fabric and OneLake. Pick any metric on a Power BI dashboard. With Atlan, you can trace it from the source column, through the lakehouse and the semantic model, to the visual someone is about to make a decision on. 👇 @msPartner
1
5
12
918
What's new: → Atlan's Fabric connector is generally available (v2.0) → Column-level lineage from source to Power BI visual → Atlan context in OneLake as open Iceberg tables, queryable from Fabric with no copies
1
3
117
Atlan retweeted
Added two models to Decision Bench—949 shared text cases: Sage: 92.1% accuracy · 733 ms median Tev1 4B: 85.4% · 387 ms median Results: decisionbench.ai Thanks @levantolabs @bigironchris @marco_derossi @togethercompute @nutlope
Made decisionbench.ai to test because I wanted to see how well Jev actually performs compared to other models. Tested 12 models across 35 tasks. Jev did really well for how little it costs. Wrote up what I found 👇
Article

Putting Jev to the test with Decision Bench

12 AI models tested on 1,071 cases across 35 tasks and 11 domains. With all the hype around Jev and the experiments people were sharing, I wanted to see how useful it actually was. I was interested in

4
9
215
The scene is set folks! Bringing together folks to make sense of Multiplayer AI, IRL in San Francisco on the October 1st :) Excited to announce and celebrate the partners making this happen @AtlanHQ , @hydra_db and @UseCorgi. Bring your questions, curiosities, demos, learn and share notes.
4
4
11
6,079
Read this as an architecture diagram, not a leaderboard!! @rohanatlan built decisionbench.ai/ to test Jev on the small decisions we keep handing to models: 1,071 real cases across 35 tasks. On the same text-only cases, Jev and Gemini 3.5 Flash were basically tied on accuracy (93.2% vs 93.5%)....but Jev was about 8x faster and about 70x cheaper in his setup! Next thing I'd want to know is whether Jev knows when it's wrong, aka how reliable is its confidence...
Made decisionbench.ai to test because I wanted to see how well Jev actually performs compared to other models. Tested 12 models across 35 tasks. Jev did really well for how little it costs. Wrote up what I found 👇
Article

Putting Jev to the test with Decision Bench

12 AI models tested on 1,071 cases across 35 tasks and 11 domains. With all the hype around Jev and the experiments people were sharing, I wanted to see how useful it actually was. I was interested in

1
5
232
How well does Jev hold up on the decisions a company actually makes? Our team tested it on 1,071 real cases across 11 domains, from finance to legal to support. The result: 93.8% accuracy, under half a second per answer, about four cents per 1,000.
Made decisionbench.ai to test because I wanted to see how well Jev actually performs compared to other models. Tested 12 models across 35 tasks. Jev did really well for how little it costs. Wrote up what I found 👇
Article

Putting Jev to the test with Decision Bench

12 AI models tested on 1,071 cases across 35 tasks and 11 domains. With all the hype around Jev and the experiments people were sharing, I wanted to see how useful it actually was. I was interested in

1
5
624
Atlan retweeted
The most underrated value of AI — in software, we had to guess what the user wanted. With conversational AI — they tell us. For a user obsessed company, this has been a gold mine. @rishigb expands how we have turned traces to empathy at scale — and how it is helping us compound user value.
Everyone is talking about evals and judges as possibilities. We put one in production. It scores every MCP outcome for real user value and turns the failures into engineering work. I built it from 1,000 conversations read by hand. Here is the whole thing.
Article

Empathy at Scale: How an LLM Judge Improves Our Atlan MCP Server

TL;DR: An LLM judge scores every Atlan MCP outcome for real user value, and a loop turns the failures into shipped engineering fixes. I built it from 1,000 conversations read by hand. At Atlan, we are

1
1
11
2,434
For most teams, an eval is a report card. Ours is a control loop. 👇
Everyone is talking about evals and judges as possibilities. We put one in production. It scores every MCP outcome for real user value and turns the failures into engineering work. I built it from 1,000 conversations read by hand. Here is the whole thing.
Article

Empathy at Scale: How an LLM Judge Improves Our Atlan MCP Server

TL;DR: An LLM judge scores every Atlan MCP outcome for real user value, and a loop turns the failures into shipped engineering fixes. I built it from 1,000 conversations read by hand. At Atlan, we are

3
451
If any of this is what you're building, or just what you can't stop thinking about, come. Oct 1, SF. A room of people figuring it out, real prototypes, hard questions. Apply to demo or RSVP here: luma.com/fj1r2uu7
1
3
392