Open-Source AI Observability and Evaluation

Notebook or Container
your agent can think fast and slow. kahneman's "thinking, fast and slow" describes two modes: system one: fast, cheap, almost automatic (jev) system two: slow, deliberate, analytical (llms)
1
1
3
263
now you can compose both. a small model handles routing and guardrails, and the llm steps in only when needed. openinference's new decision span captures the question, choice, confidence, and probabilities, so system one calls show up right next to the reasoning they gate.
1
52
below: a support agent traced in phoenix. jev triages the ticket in 237ms and judges whether the reply is ready to send in 154ms, bookending two ~1s openai calls. 2.2s end to end, under a cent. ships today in our typescript and python auto-instrumentors.
44
Benchmarking AI agents & tool use with Harbor + Arize Phoenix. Join us Oct 8 at 11am PT / 2pm ET to learn how to compare agents or models over the same task set, separate behavioral scores from infrastructure failures, and inspect ATIF traces. luma.com/arizeai-benchmarkin…
2
3
171
Not all tokens cost the same. Cache writes have an upfront cost, but they save you money on subsequent turns. Premature or inefficient compaction can cause cache misses that add both latency and cost. In Phoenix, you can search LLM calls across a turn and see how context evolves.
2
1
5
332
Webinar coming up: benchmarking AI agents & tool use with Harbor and Arize Phoenix. Learn how to compare agents and models on the same tasks, diagnose regressions, and inspect scores, errors, and traces. Oct 8 · 11am PT / 2pm ET Save your spot: luma.com/arizeai-benchmarkin…
7
1,830
When conducting experiments, it might be important to keep track of the data being fed into the agent using a Git like version control system - especially if you are collaborating with others. Phoenix now has bulk editing for datasets so that you can track the dataset patches.
1
3
220
TypeSafe came out of stealth this week with Jev, the first System One Model: a frontier model that never generates prose. Unstructured state in, typed probabilistic decisions out.
1
2
4
760
Here it is blocking a support draft that promised free shipping for life, at confidence 1.0. Python: pypi.org/project/openinferen… JS: npmjs.com/package/@arizeai/o…
1
59
“Did the agent actually do what the user asked?” 👀 That is the question behind many agent reviews. A trace can look good but still miss part of the request. The retrieved context can be real but unhelpful. The response can add facts the conversation never supported.
1
2
189
🛡️ Toxicity checks harmful or abusive text. 🔒 PII detection checks conversation records for personal data. The pages show the needed input fields and examples in Python and TypeScript. They also explain the score output so you know what the result means.
1
1
46
For teams reviewing traces or comparing experiments, this is a practical starting point before writing custom judges. Start here: arize.com/docs/phoenix/evalu…
23
With the new OpenInference instrumentation for ADK Java, every agent invocation, LLM call, and tool execution shows up as a trace in Phoenix. Same experience the Python ADK folks have had, now for the JVM.
1
45
If you're building agents in Java, we'd love to hear how this fits your setup and what's missing.
26