Golang/observability hacker changing the world @typesafeai. alum @docker @honeycombio @Bauplan_labs

San Francisco, CA
still seems expensive to me
If you want to understand how fast AI is improving, look at intelligence per dollar. 1.5 years ago: o1 Pro: $150 / $600 per M tokens Today: GLM-5.3 Flash: $0.15 / $0.50 That’s a ~1000x collapse in price in under 1.5 years, while GLM-5.3 Flash is more intelligent than o1 Pro.
5
39
18,031
one thing good about writing shitty code the first time is you can come in and crush -30%, -50% -90% allocs or cpu cycles
80
hyd today i'm making jev even faster lg
6
26
1,297
very cool. when we first started making demos i made an @opencode fork that would dynamically choose if you typed bash command, python or regular chat on enter and run it. the fluidity of not having to type !, etc. felt really nice. tools should adapt to us!
Jev is now in the AI SDK for Python. To test it, we ran two experiments: ▪︎ Detecting Python vs. English as text is typed ▪︎ Writing Python, one decision at a time 𝚞𝚟 𝚊𝚍𝚍 𝚊𝚒 vercel.com/blog/jev-for-pyth…
2
11
1,208
“he’s basically jev’s only salesperson”
Apparently the biggest flex at the club in sf is a full Google Calendar
1
1
25
1,365
Nathan LeClaire retweeted
Replying to @dotpem @danielgafni
not rip at all, the edge appliance meta is coming. quantize it down to run on a mini pc or homelab box and you don't even need the datacenter. zero network latency beats everything
1
1
107
Nathan LeClaire retweeted
Replying to @dotpem
1
2
97
quem quer um jev em latam? (não tô prometando, só curioso)
Replying to @dotpem @danielgafni
latam rtt is the real barrier: 140ms roundtrip to us-east kills the whole advantage of a 15ms decision model. dropping an inference node in são paulo puts local routing under 10ms for real-time messaging agents. devs here would run it non-stop
4
15
811
same
If you did something cool with Jev and codemode I want to know!
1
477
you also get to meet @hmartenjoyer which is probably the best part
Got Jev? It's easier than you think!
1
1
7
707
my weekend plans are to work on jev so you degenerates can implement systemone routers that work even faster... and support more things ;)
CANCEL your weekend plans. You NEED to: • Replace fragile chat loops with Temporal state machines for crash-resilient agent workflows • Build a dynamic context assembler that strictly budgets tokens between memory, tools and docs per request • Implement a System-One router to handle 95% of triage in milliseconds, escalating only edge cases to reasoning models • Write a raw Model Context Protocol (MCP) server from scratch to expose your database securely to external agents • Build trajectory grading in CI that blocks PRs if an agent's tool-call sequence deviates from the golden path • Add an inter-agent sanitization proxy so Agent A's output can't execute a prompt injection on Agent B • Implement a cost kill-switch that halts any agent loop if projected token spend exceeds $0.50 per query • Build a 3-tier memory engine (Working, Episodic, Semantic) with automated eviction and compression policies • Set up KV-cache prefix sharing at the proxy layer to slash Time-To-First-Token by 80% for identical system prompts • Route 5% of production traffic to a new model silently, compare the trajectories and auto-generate a diff report • Build an API-to-Browser fallback where the agent spawns a headless browser if the REST API 404s • Automate a DPO pipeline: user thumbs-down → auto-format to preference pairs → queue a nightly LoRA fine-tune • Implement a PII unmasking proxy that swaps sensitive data for UUIDs before the LLM sees it and restores it post-generation • Set up a speculative decoding pipeline: local 2B model drafts tokens, cloud 70B verifies them, cutting latency 60% • Build a 4-tier graceful degradation chain: Frontier API → Mid-tier → Local Quantized → Semantic Cache • Add checkpointed human-in-the-loop approval gates wired directly into Slack for high-stakes agent actions • Trace every LLM hop with OpenTelemetry-style spans capturing exact token counts, latency and tool arguments • Fuzz your own multi-agent swarm with malformed tool outputs and context overflow to test self-healing recovery You have way too much to do. Bookmark & Repost.
10
40
2,990
Nathan LeClaire retweeted
Decision models in llama.cpp are now available The `/v1/systemone` endpoint is available in the latest llama builds. Use it to do Jev-style inference locally, efficiently and privately. Multiple open models are supported with more to come. huggingface.co/blog/ggml-org…
89
354
2,155
87,805
the older i get the more i begrudgingly accept that the capn proto guy was probably right
2
1
6
709
btrfs? we're about due for it to be production ready, no?
I'm speedrunning my filesystem learnings. Turns out, brtfs cow-feature is GREAT for worktrees and terrible for sqlite. Next OC update knows that and - if you did sth silly like me - will migrate your database to a NOCOW spot.
2
706
Nathan LeClaire retweeted
This office hours recording from @DSPyOSS is a gold mine of excellent Jev use-cases and explainers!
Today at 2pm PT, @isaacbmiller1 and I will be walking through @DSPyOSS's Jev/System One implementation & the ReAnchor optimizer. Will also sharing some design patterns for for selective compaction, tool approvals, and subagent delegation using Jev. streamyard.com/watch/sbMTpJA…
4
11
155
23,210
jev secretly uses human optic nerves in the datacenter confirmed?
“For you to see”?? See with Jev?! Multimodal confirmed??!!
14
691
do i know anybody that works in ai model containment
12
22
4,290
we meet again cgo
1
1
433
Nathan LeClaire retweeted
me to our GTM guy: do enterprises want benchmarks him: no, the fast ones have already benchmarked their internal use cases. the slow ones are copying the fast ones. he's onboarded a quarter of the fortune 500 already 🥵
52
33
1,003
82,964