Writing about and working on AI, DSPy, geo, and data. Trying to prevent prompt debt and help others build maintainable intelligent applications.

Bay Area
A quick example of a model router, using DSPy and Jev, based on your initial prompt. Sketched this out in ~15 min, by request, and am surprised at how well it works! gist.github.com/dbreunig/949…
10
4
85
11,706
No OS has had a more cohesive vision than the Wii.
3
9
748
Reasoning is so useful for AI engineering, even summarized reasoning. The fact that Opus and Sonnet 5.5 both shut down when asked to summarize why they took an action is causing people to choose other models.
The fact that I can't ask Claude to "think out loud" without getting shut down for so-called "reasoning extraction" is just embarrassing. It's clearly not stopping the real distillers and it makes it impossible to use this product because it just spends a minute and a half thinking and then lies to me about what it thought about. Sorry, but the reasoning is the good part! I don't need a crappy overcooked deliverable. I need a partner in thinking
4
3
24
1,641
Drew Breunig retweeted
This office hours recording from @DSPyOSS is a gold mine of excellent Jev use-cases and explainers!
Today at 2pm PT, @isaacbmiller1 and I will be walking through @DSPyOSS's Jev/System One implementation & the ReAnchor optimizer. Will also sharing some design patterns for for selective compaction, tool approvals, and subagent delegation using Jev. streamyard.com/watch/sbMTpJA…
4
11
156
23,453
Here's a repo demoing these compound AI patterns with DSPy + Jev! github.com/cmpnd-ai/dspy-sys…
Here's the post. Really enjoyed these thoughts. It's a great exercise to reconsider the design choices KV caches make for us.
2
11
105
8,547
Drew Breunig retweeted
the @DSPyOSS town hall about dspy+jev was really good. here are some of my favorite highlights
2
11
70
16,740
Also @hammer_mt will drop by!
Today at 2pm PT, @isaacbmiller1 and I will be walking through @DSPyOSS's Jev/System One implementation & the ReAnchor optimizer. Will also sharing some design patterns for for selective compaction, tool approvals, and subagent delegation using Jev. streamyard.com/watch/sbMTpJA…
2
7
1,033
Today at 2pm PT, @isaacbmiller1 and I will be walking through @DSPyOSS's Jev/System One implementation & the ReAnchor optimizer. Will also sharing some design patterns for for selective compaction, tool approvals, and subagent delegation using Jev. streamyard.com/watch/sbMTpJA…
2
4
42
23,397
Someone please exfiltrate the system instructions for america.gov
2
7
1,326
We’re considering hosting a short, casual @DSPyOSS webinar next week to run through Jev support and design patterns. Let me know if you’re interested, and what you’d like to see covered.
14
5
96
3,565
This Wednesday at 2PM Pacific, @isaacbmiller1 and I will be walking through DSPy's Jev integration, the ReAnchor optimizer, and discussing a few design patterns. (Including some detailed by @CompleteSkeptic in his recent post on agents) Join us! streamyard.com/watch/sbMTpJA…
1
3
37
5,387
Here's the post. Really enjoyed these thoughts. It's a great exercise to reconsider the design choices KV caches make for us.
sharing some notes on typesafe 🤝 coding agents: docs.google.com/document/d/1… we likely will never have time (ever again) to play ourselves, but hope the that the community goes WILD (and makes me look like a naive idiot)
1
20
9,613
Drew Breunig retweeted
I wrote a short blog on my experience testing Jev on reasoning-intensive regression problems. The results are within expectations (Jev is focused on classification after all), but the model has great attributes and is wicked fast! Read about it here: dianetc.com/musings/jev/
2
12
103
21,288
tbf, in the Claude Code leak we saw Anthropic was using regex for swears to flag threads for QA review.
Jev agrees that I'm vibing more with Opus.
2
6
1,630
Just released dspy-monty-interpretter, bringing monty 1.0 to all your dspy.RLMs! github.com/dbreunig/dspy-mon…
Fuck it, still early but here goes ... We've just released Monty v1 - a Python sandbox that starts in 1 millisecond, not 1.5 seconds. I just ran 10k sandboxed scripts in 674ms, something that would take a cloud sandbox > 3 hours. This removes the biggest drawback of letting agents write code. The future is fast. Even better, it's open source, you can install it from PyPI, npm or Crates now. Serviced platform coming soon. Please get in touch if you want to be a design partner! Who should try it? ⚡ if you care about startup time, use Monty ⚡ if you care about long-lived sessions, use Monty - Monty can be dumped and resumed at any external function call ⚡ if you care about accessing functions in the agent/host, use Monty - Monty makes it trivial to expose local functions into the sandbox ⚡ if you care about scale, use Monty - Monty workers use as little as 2MB of memory, meaning you can run thousands of concurrent sandboxes on a single machine ⚡ if you care about security, use Monty - we've run 3 rounds of bounty program and thousands of researchers have tried to break into our sandbox, meaning it should be secure to run untrusted code Who should avoid it? 🚫 if you like to take a coffee break while waiting for sandboxes to start, DO NOT use Monty 🚫 if you enjoy the challenge of routing API requests from sandboxes through your corporate network to access state in your agent without exposing secrets to the sandbox, DO NOT use Monty 🚫 if your agent really needs to install packages from PyPI, Monty won't help you yet (spoiler: it probably doesn't) pydantic.dev/docs/monty/get-…
5
4
90
7,153
The most recent batch of models start to respond in a limited and localized vocabulary as you go over ~200k context. Unique terms are adopted, more like pointers (symbolic) than jargon (language). At least it feels that way to me. I end up spending significant effort getting them to “decompile” this dialect when writing reports, or even product copy.
I bet we figure out that neural networks can be decomposed into evolved symbolic systems of sorts and that it will be achieved before capital S Superintelligence but everyone has to lock in
1
21
1,621
So often, the answer is: an RLM with tools.
7
1
66
3,792
What’s your favorite skill you’ve made?
3
7
867