JEV: Your AI Agent Is Burning Money While You Sleep
Imagine your agent at 2:46 a.m. The patch is ready. Somewhere inside the loop, a tiny committee is still asking whether to inspect another file, run another check or announce victory. Every committee meeting wakes the same expensive model. In this imaginary control room, the rocket is on the launchpad and the main engine is being used to ring the doorbell.
Open a real trace and look for the equivalent. A model call that ends in one label may still carry thousands of input tokens. Jev gives those bounded judgments a separate route: send the state and explicit questions, receive typed answers, then let the host decide what happens. Claude keeps the planning and code generation. The expensive engine gets a more selective ignition system.
The first places to look are brutally ordinary: the check before an action, the check before stopping, and the check before starting another session. Move one suitable decision, measure the full run, and see whether the bill changes. In the four-agent example below, decision inference alone costs $1,415.04 through Reference A or $5.59 through Jev. The rest of the agent still has a meter.
Find the Tiny Questions Running Up the Bill
Your next optimization target might be a question so boring that nobody thought to price it. Which worker next? Is this request a bug? Does this evidence support completion? Open the transcript and mark what each call actually produced. Keep code generation and substantial planning separate from selecting an item on a menu.
Look for a small answer space, a clear rubric and enough repetition to hurt. The host should know every action it offers. Collect missing facts before asking the model to judge them. Use an ordinary parser wherever it can determine the answer exactly. There is no prize for adding inference to a comparison your code can already make.
| Decision | Possible output | What remains outside Jev |
| --- | --- | --- |
| Choose the next worker | research / implement / verify / review | Starting the selected worker |
| Assess a proposed action | routine / sensitive / uncertain | Permissions and execution |
| Review completion evidence | Probability that a requirement is met | Tests and direct outcome checks |
| Triage incoming work | bug / question / feature / other | Queueing and starting sessions |
A fraction-of-a-cent router can still send your expensive worker on a useless expedition. Count the wrong turns, the retries and the human cleanup. Move one branch first and compare the whole completed task. Otherwise, you may celebrate the cheap ticket while quietly buying three connecting flights.
Stop Asking for Essays When You Need a Switch
Jev reads text or structured text such as JSON. Choice selects among options you define. Score evaluates against ordered levels. Noul returns a probability for a yes-or-no question. Choice and Score also expose confidence derived from their probability distributions; Noul does not have that extra field. These are answers your program can consume without digging a label out of a miniature essay.
Think of the schema as a control panel. The buttons have names, and the wires go somewhere specific. That still leaves the possibility of pressing the wrong button. Keep an other or review route in the menu, explain each category, and evaluate the actual choices. A perfectly formatted wrong answer is still a wrong answer, just dressed for the meeting.
Install typesafe-sdk, set TYPESAFE_API_KEY in the environment, and run one bounded request. The example below asks two independent questions about the same snapshot. It prints a recommendation and performs no workspace action. Start there: a result you can inspect before you connect it to anything that moves.
from typesafe_sdk import Choice, Noul, TypeSafeClient
state = {
"goal": "Fix the CSV parser and document the new flag.",
"patch": "Parser changed; the new flag is implemented.",
"checks": "Unit tests passed. README still needs review.",
}
questions = {
"next_step": Choice(
instructions="Choose the next step using the supplied evidence.",
criteria={
"inspect": "More repository evidence is needed.",
"implement": "A required code change remains.",
"verify": "The change needs checks or documentation review.",
"review": "The request is unclear or outside these options.",
},
),
"documentation_missing": Noul(
instructions="Does the evidence indicate unfinished documentation?"
),
}
with TypeSafeClient(model="jev-1.13.0") as client:
result = client.system_one(state=state, questions=questions)
route = result.choices["next_step"]
print(route.choice, route.confidence)
print(result.nouls["documentation_missing"].noul)
Put the meaning in instructions and criteria. Naming a question safe_to_run does not magically upload your permission policy into the model. Keep the model version, question definitions and thresholds together. When something changes, you need to replay old cases and see which decisions changed with it.
Stop Paying Four Times for the Same Snapshot
Picture a fictional spaceport with four inspection desks. Each wants the same flight manifest, but asks a different question. Is the route clear? Is documentation missing? Is the fuel evidence sufficient? Does the mission need review? Make the pilot submit a fresh copy at every desk and you pay to transmit the manifest four times before anyone has moved an inch.
Jev can evaluate multiple questions in one request in parallel against shared state. The questions still add input tokens; the document does not need to make a separate trip for each one. Bundle independent judgments, then combine the answers in code. The saving comes from changing what you send and how often you send it.
Use a budget of 3,000 shared state tokens and 150 tokens per question, including its instructions and schema. Four separate requests consume 12,600 input tokens. One request carrying all four questions consumes 3,600. That is 71.4% less billed input, or a 3.5x reduction in this example. This is token arithmetic, not a measured latency result; the question budget must cover every billed part of each question.
Keep the dependency boundary real. Questions inside a batch cannot read one another's answers. If judgment two needs a search result, run the search and build a new snapshot first. Batching can remove repeated input. It cannot conjure evidence that has not arrived.
A Probability Is Not a Production Credential
The thrilling part of an autonomous demo is watching the agent act. The expensive part can be discovering what it acted on. A command needs context: working directory, target files, environment and permitted scope. Apply the exact restrictions your application can enforce, then give the classifier the relevant evidence for the judgment that remains.
A probability is not a production credential. Jev can flag an ambiguous action for review; the host's permission policy controls execution. Confidence does not create access rights. Even running a test suite can mutate files or touch services. Give the model a defined job inside the boundary, and keep the keys with the code that enforces it.
Claude Code exposes PreToolUse before a tool executes. A hook can request confirmation with permissionDecision set to ask, as below. When the hook has no additional restriction, leave the existing permission flow in place. Turning every confident label into an automatic allow would give a classification result a job it cannot justify.
{
"hookSpecificOutput": {
"hookEventName": "PreToolUse",
"permissionDecision": "ask",
"permissionDecisionReason":
"The proposed action needs review under this project's policy."
}
}
Decide what happens when the signal disappears. A timeout, unfamiliar label or conflicting evidence should leave the action pending or send it for review. Log the proposed action, model result, policy result and actual outcome separately. When the system surprises you, that record is the difference between finding the fault and staring at a beautifully confident number.
Make Done Prove Itself
The word done is wonderfully cheap to generate. Evidence is harder. Start with the requested outcome: if tests must pass and a file must exist, check both directly. Then give Jev the part that needs interpretation: does the evidence cover the remaining requirements, or did the worker merely write a convincing farewell?
Send the test output, relevant diffs and acceptance criteria. In an imagined bad run, the tests glow green while the README still teaches last week's command. The agent delivers a victory speech; the next user gets an error. This is exactly the kind of gap a completion review should be designed to examine.
A Stop hook can ask Claude to continue. Check stop_hook_active and use a bounded retry policy. If the cap is reached while requirements remain unmet, mark the task incomplete or for review. A stopped process and a completed job are different events. Your reporting should preserve that distinction.
Log every forced continuation. An evaluator that keeps rejecting correct work becomes its own little money furnace. Measure correct completion decisions on your tasks, including cases where the right answer was to let the worker stop. The loop needs a reliable exit as much as it needs a reliable alarm.
Stop Waking an Expensive Brain for Nothing
There is an even earlier place to look at the meter: before the worker starts. Read newly changed issues, remove duplicates in code, then classify what remains. A reproducible defect can enter the coding queue. A usage question can go to support. An unclear request can wait for review. Let the queue earn the expensive session it is about to consume.
Use an issue ID plus an update marker so the same task cannot keep returning in a different hat. Reserve the item before starting a worker, store the session reference, and update its status when the worker returns. A cheap classification is a small win. Preventing a duplicate coding session can affect a much larger piece of the run.
Do the first rollout beside the existing queue. Compare labels before allowing them to suppress sessions, and keep a fallback for missed work. A filter that makes the dashboard quieter by hiding real bugs has found a very creative way to lose money. Measure useful work reaching completion, not just fewer workers waking up.
Four Agents Can Build a Four-Figure Decision Bill
Now put a small agent team on an invoice. Four agents each run 120 cycles per night, with four independent checks per cycle, across 22 working nights. That creates 1,920 decisions a night and 42,240 for the month. Use the same 3,000-token snapshot and 150-token question budget from the batching example: 3,150 billed input tokens per separate call. The generative references produce 40 output tokens per check. A quiet background task is starting to look like a payroll item.
|Decision model | Input / output per million tokens | Per night |22 working nights |
| --- | --- | --- | --- |
| Reference A | $10 / $50 | $64.32 | $1,415.04 |
| Reference B | $3 / $15 | $19.296 | $424.51 |
| Jev | $0.042 / $0 | $0.254016 | $5.59 |
At the comparison rates shown, Reference A reaches $1,415.04; Reference B reaches $424.51; Jev reaches $5.59. Reference A and B are illustrative rates, not claims about your subscription. Jev's listed price is $0.042 per million input tokens with no output-token charge. Totals are calculated before rounding. Caching, discounts, tools and code-writing calls are excluded. This invoice covers the decision workload alone.
Now batch the same four independent checks. The team needs 10,560 requests over those 22 nights, each carrying 3,600 billed input tokens. At the listed Jev rate, that is $1.596672, or $1.60 rounded, against $5.588352 for separate Jev calls. The monthly input volume falls from 133.056 million to 38.016 million tokens. You changed how often the snapshot crosses the wire. Count every question token too: good numbers should survive someone opening a calculator.
Your actual bill still has to pass the same test. A subscription does not refund cash every time you avoid a token. Measure the usage budget and completed work under that plan. For API billing, count paid usage across the entire run, including the mistakes and retries. Put that number beside the old one before declaring victory.
Build a Museum of Wrong Answers
The clean demo is the easy part. Give the system your awkward cases: a stale tool menu, a vague request, missing evidence, a valid option that was still the wrong choice. Begin in shadow mode, where Jev records recommendations while the existing system stays in charge. You want to discover the strange failures while they are still entries in a log.
Track decision accuracy, fallback rate, total latency and cost per completed task. Treat confidence as a model output and correctness as something you observe. Choose thresholds using labeled cases that resemble production, then keep a separate evaluation set for the final comparison. An impressive number has to earn its place in the dispatch code.
Jev's documented limitations include arithmetic, date comparisons, irrelevant context and adversarial instructions inside the state. Keep exact calculations in code, send the evidence that matters, and test your boundary cases. Stuffing the entire repository into every request gives the model more text; it does not automatically give the decision more useful context.
Replay the same cases whenever you change the rubric, model version or candidate list. Keep the failures that taught you something. Logs give you material for improving the surrounding system; they do not secretly turn it into a self-training intelligence. The improvement comes from changes you can identify and measure.
Take One Decision Off the Meter Tonight
The opportunity is sitting inside a trace you probably already have. One repeated judgment. One bounded answer space. One place where a costly worker is doing a smaller job. Your first result might be fewer unnecessary sessions, better completion checks or cheaper routing. Keep the baseline and find out which improvement actually happened.
For a team, this can become a concrete service: inspect the run, identify a decision worth moving, install the evaluation path and report the effect on completed work. Sell the measured outcome and maintain the rules around it. A working branch with a defensible result is easier to assess than another grand announcement about the future of agents.
Save this and open one overnight trace tonight. Pick the decision that repeats. Define its options, write the fallback and run Jev beside the current path. Compare the results before moving control. Your agent may already have enough intelligence to do the job. Find out how much you are paying to make it keep asking what happens next.

