The AI engineering platform for teams shipping reliable AI agents and LLM applications. Also home to @ArizePhoenix.

San Francisco, CA
Pinned Tweet
Jev-as-a-Judge is live in Arize AX. Ask most LLM judges a yes/no question and you get a paragraph back. Then you parse it to figure out the answer. Jev returns typed outputs instead: booleans, choices, or rubric scores, with confidence on classifications. You can also evaluate multiple questions in a single call. Docs --> arize.com/docs/ax/evaluate/j…
9
3
20
881
More prompt caching ≠ lower costs. We benchmarked DeepSeek, GLM, GPT, and Claude across 400 multi-turn agent runs. Claude reused 89.8% of prompt tokens but still had the highest estimated cost. Here's what the traces revealed 👇 arize.com/blog/prompt-cachin…
3
7
399
We built long-term memory for Alyx with retrieval and a knowledge graph, then shipped an 8,000-character file that loads on every request instead. Here's the architecture, tradeoffs, and evals behind that decision (as well as a deep dive into how we approached building long-term memory for our agent Alyx): arize.com/blog/alyx-agent-lo…
3
1
5
453
New ways to work with Arize AX: we now have a hosted MCP server that you can use directly with your coding agent. Learn more: arize.com/blog/extending-you…
earlier this year our CPO @aparnadhinak posted "MCP = context tax. the future is CLI and files." last week we shipped a hosted MCP server for Arize AX 😅 both decisions were right; harnesses just caught up with tool search and code execution. once that token tax stopped mattering, the real question became where your agent was and who was asking for data wrote up how we think about MCP vs CLI vs Skills, with receipts from both the @OpenAI + @AnthropicAI folks who've been discussing this on the timeline 👇🏽 arize.com/blog/extending-you…
1
1
248
Arize AI retweeted
your agent can think fast and slow. kahneman's "thinking, fast and slow" describes two modes: system one: fast, cheap, almost automatic (jev) system two: slow, deliberate, analytical (llms)
1
1
4
274
September ships: → Jev-as-a-Judge for high-volume structured evals → Agent-as-a-Judge now available on every plan → Alyx long-term memory + experience-aware help → Full session + trace annotation queues → Remote evaluator runs you can inspect as traces → GPT reasoning in Alyx → New model support in the Playground + evaluators And lots of improvements across tracing, evals, annotations, APIs, SDKs, and Signal. Full September changelog: arize.com/docs/ax/release-no…
1
1
8
504
Alyx now remembers what you taught it yesterday. Tell it where ground truth lives, what success means for an eval, or why a prompt changed. Open a new session tomorrow and Alyx can use that context as you keep working. Long-term memory is now live in Arize AX: arize.com/blog/alyx-arize-ax…
3
6
367
We built it around one deliberately small constraint: an 8,000-character memory file per user and space. Alyx updates memory in the background using targeted inserts, replacements, and deletes, which keeps changes inspectable and limits drift. Engineering deep dive: arize.com/blog/alyx-agent-lo…
1
120
Building agent memory means solving a few deceptively hard problems: - What should survive? - When should it be used? - What happens when a fact changes? - What should the agent forget? We wrote the guide we wish we had before building ours: arize.com/resources/long-ter…
1
71
A great deep dive into how we built long-term memory in our AI engineering agent Alyx from the engineer behind it. 👇 Get the full write up: arize.com/blog/alyx-agent-lo…
Memory or context beyond sessions is something all agents and harness providers are still trying to perfect. There’s many different solutions, and it depends on what your agent can do. I wrote this blog post on how I built the most effective long-term memory system for Alyx, an AI engineering agent that analyzes your agent traces, creates evals, and runs experiments for you. I broke down how I chose my system, and why it’s similar to what OpenAI and Anthropic are converging to as well. arize.com/?p=32589&preview=1…
1
2
501
ICYMI: Moving past the hype with TypeSafe AI's Jev model, how we used a managed agent to find 43 unnecessary tool calls, a conversation with Daytona's founder Ivan Burazin, and what's new in Arize AX.
Article

What Jev is teaching us about better evals

We've been busy testing TypeSafe’s Jev across tens of thousands of evaluations. One result stood out: Jev’s confidence was a much stronger signal of its own mistakes than repeatedly running an LLM

17
1
19
629
One Jev call gave us a better error signal than rerunning an LLM judge 10 times. Jev’s probability reached 0.95 ROC AUC for identifying incorrect judgments. Disagreement across 10 repeated LLM judge calls reached 0.62-0.84. We tested this across 36,190 judgments and 10 Phoenix evaluators. More in 🧵 and the full benchmark's here: arize.com/blog/jev-llm-judge…
2
8
380
The confidence score was useful on its own. When Jev reported ≥99% probability, all 3,561 judgments in that bucket were correct. Below 70%, accuracy fell to 73%. That gave us a signal for deciding which evals deserve another look.
2
2
154
We walked away with a practical pattern for production evals: keep the probability returned with every judgment, then send lower-confidence cases to human review. That can give you an uncertainty signal from one call instead of estimating it through repeated judge runs. Full benchmark: arize.com/blog/jev-llm-judge…
1
1
92
Arize AI retweeted
Jev-as-a-Judge is live in Arize AX. Ask most LLM judges a yes/no question and you get a paragraph back. Then you parse it to figure out the answer. Jev returns typed outputs instead: booleans, choices, or rubric scores, with confidence on classifications. You can also evaluate multiple questions in a single call. Docs --> arize.com/docs/ax/evaluate/j…
9
3
20
881
Arize AI retweeted
@seldo compared Jev, Claude Opus 5, and GPT-5.6 Terra with 23,325 judgments on accuracy, cost, latency, and calibration. We found that Jev matched Claude Opus 5 at 87% hallucination-detection accuracy while running 23x faster and at roughly 1/300 the cost. arize.com/blog/jev-llm-judge…
3
3
17
1,190
When using an LLM-as-a-judge, or playing with prompts, you need access to the models you want at high speed, including open source, and your fine-tuned models. Arize AX now has support for @FireworksAI_HQ as an AI provider, so you can use all your favorite models. arize.com/docs/ax/security-a…
1
8
420
Every session with an AI assistant starts from zero. You have to re-explain the goal, the error column, and the spike that turned out to be test traffic. Alyx now remembers your goals, the signals you watch, what you already ruled out, and more. Long-term memory docs --> arize.com/docs/ax/alyx#long-… Experience awareness docs --> arize.com/docs/ax/alyx#exper…
1
2
309
But with Jev’s default threshold, it looked 7 points worse. At the default 0.5 cutoff, Jev scored 76%. We tuned the threshold on human-labeled validation data, then tested it on held-out examples. At 0.8, Jev reached 87%, with the same false-alarm and miss rates as Opus 5.
1
115