“Study hard what interests you the most in the most undisciplined, irreverent and original manner possible.” ― Richard Feynmann I am currently studying: ralph wiggum, rlms, dspy, gastown & related agent orchestrators, agent-native apps and orgs How about you?
5
49
7,436
Raymond Weitekamp retweeted
Today we're unveiling Trillium Labs @trillium_labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next. We're built around the theory of change that you need more eyes to solve hard technical problems. We have faith in the scientific methods and communities that humanity has built, and worry that AI is becoming too closed to utilize them. Trilliums are wildflowers that bloom briefly in the spring, before the forest canopies fill out. Though they are small, they lay the foundation for the cycles of growth and nourishment through the rest of the year. At Trillium Labs, the recipes will be the slow nutrients for the seasons and the model releases will be the blooms. Building an institution dedicated to this is needed because, much as nature’s trilliums are slow to expand and grow, the open-ecosystem needs time and dedicated resources to catch up. I co-founded with with a long-time friend and collaborator Tom Zick (@thesezickbeats). We're hiring (full time + student collabs/interns), we're fundraising, and we're looking for compute. Please get in touch if you're interested in helping out. Offices based in the Bay Area and Cambridge MA, remote okay. I’m in the Bay Area until for The Curve and COLM to connect with people who are interested. We’re thankful to have initial support from Halcyon Futures and Schmidt Sciences with more funding en route to enable our ambitions of scaling. Our advisors @Thom_Wolf, @HannaHajishirzi, @gneubig and @ctnzr have been instrumental to building the ecosystem that exists today, and I’m stoked to get to keep working with them.
238
275
2,950
192,781
this looks awesome
introducing 𝚌𝚕𝚎𝚏: our first models trained by @cloudflare's workers ai team. today, we're releasing two fast and accurate decision models that top the benchmarks for quality and latency. use them hosted on workers ai or grab the weights from @huggingface, because we open-sourced it too. blog.cloudflare.com/clef-dec…
4
404
i am so stoked for this! (plus a little jev sprinkled on top)
For years, maybe a decade, I've wanted to build a writing tool that just did a few things I couldn't find anywhere else. A system that mimicked how I think when I write. A system that provided clear affordances, features, and places to do the things I wanted to do. A way to play with my writing as I'm writing it. A way to settle into just the right way to say something. It never made sense to make it a 37signals product because it's not commercially viable, nor would it be worth pulling people off their other work to hack on this. Plus, it would have taken months the old way. So this weekend I just made it myself, with Claude's help. It's called Write_On and the short video walks you through it. Essentially it's "alternative control" at the word, sentence, and paragraph level, plus a way to dim stuff back, and stash stuff near by. Not version control, but alternative control. You'll see what I mean when you watch the video below. So here it is. Write_On.
2
581
Raymond Weitekamp retweeted
i wanted to voice dictate without bothering everyone so I made this lip reading app lol github repo in the thread
289
177
2,871
212,149
still need to read the paper, but the naming is genius... RLTL;DR
We let an agent self-improve by exploring and internalizing ”on this sort of task, keep this sort of thing in mind“ bits of self-feedback. It learned to solve tasks where the original policy failed 128 attempts in a row and GRPO flatlined. Paper: arxiv.org/abs/2609.37633 🧵1/8
3
115
33,706
Raymond Weitekamp retweeted
The latest Interconnects plot - showing the exponential growth of the open model inference economy. Plot is showing the daily tokens processed by the leading companies, based on public disclosures (for Together, Baseten, and Fireworks) and the OpenRouter API. I underestimated OpenRouter's growth.
15
11
108
15,306
Raymond Weitekamp retweeted
Today at 2pm PT, @isaacbmiller1 and I will be walking through @DSPyOSS's Jev/System One implementation & the ReAnchor optimizer. Will also sharing some design patterns for for selective compaction, tool approvals, and subagent delegation using Jev. streamyard.com/watch/sbMTpJA…
2
4
42
23,267
Raymond Weitekamp retweeted
Didn't expect this 🤯 We replaced embeddings with Jev in GPT Researcher's RAG pipeline and tested both on 28 research tasks from SimpleQA and open ended research. Jev beat embeddings on every quality measure we ran: - 59% more relevant context (73% vs 46%) - Reports preferred 15 to 3 in blind comparisons - Same cost per report GPT Researcher now runs on Jev by default, and no longer needs embeddings at all. All you need is @LangChain + @tavilyai +Jev for the perfect RAG system. Check out the repo here: github.com/assafelovic/gpt-r… Research: docs.gptr.dev/docs/gpt-resea…
95
119
1,681
271,879
80% of the time a 17 million parameter model can be finetuned to beat Jev. 40% of the time a 17 million parameter model can be finetuned, on human label, to beat Kimi-k3. 20% of the time a 17 million parameter model can be trained on Kimi-k3 synthetic data and beat Jev. An aditional 50% (so 70%), that distilled model in ball park of Jev. Sometimes, general models (like kimi and Jev), or still better because you training data does not cover well you test (prod?) data. It can take from 8 seconds to ~7 mins getting the synthetic data needed and cost between 0.77$ to ~30$. It takes between 45 seconds to 3 minutes to finetune the 17m (Ettin) model on a 3090 gpu. its 10 mins to 40 mins on cpu, and about 20 mins to 3 hrs on a samsun s21 gpu (XD just i did it!). If you have more then 500 000 inputs to classify, you should consider finetuning on synthetic data it will be both faster and cheaper then using alternatives, if you have about 150 labels per categories (human label) it will also likely be better!
17
16
192
10,228
If only I were a huge bank with millions of messages a day, I would definitely hire Maxime
Here is when and how to finetune a specialized decision model (aka classifier) that is better then Opus, faster then Jev, and runs on user's devices.
1
11
1,486
Raymond Weitekamp retweeted
Replying to @typesafeai
@typesafeai Loving behavior-driven development with Jev. github.com/lakeday-org/perch
2
1
228
38,352
Raymond Weitekamp retweeted
Introducing jevgrep - a research agent CLI powered by jev from @typesafeai that reduces your coding agent cost by 40% (verified on SWE-bench) Make sure to use the built in skill so your coding agent knows to use jg for context collection github.com/dzhng/jevgrep
193
326
4,904
445,815
Raymond Weitekamp retweeted
Fuck it, still early but here goes ... We've just released Monty v1 - a Python sandbox that starts in 1 millisecond, not 1.5 seconds. I just ran 10k sandboxed scripts in 674ms, something that would take a cloud sandbox > 3 hours. This removes the biggest drawback of letting agents write code. The future is fast. Even better, it's open source, you can install it from PyPI, npm or Crates now. Serviced platform coming soon. Please get in touch if you want to be a design partner! Who should try it? ⚡ if you care about startup time, use Monty ⚡ if you care about long-lived sessions, use Monty - Monty can be dumped and resumed at any external function call ⚡ if you care about accessing functions in the agent/host, use Monty - Monty makes it trivial to expose local functions into the sandbox ⚡ if you care about scale, use Monty - Monty workers use as little as 2MB of memory, meaning you can run thousands of concurrent sandboxes on a single machine ⚡ if you care about security, use Monty - we've run 3 rounds of bounty program and thousands of researchers have tried to break into our sandbox, meaning it should be secure to run untrusted code Who should avoid it? 🚫 if you like to take a coffee break while waiting for sandboxes to start, DO NOT use Monty 🚫 if you enjoy the challenge of routing API requests from sandboxes through your corporate network to access state in your agent without exposing secrets to the sandbox, DO NOT use Monty 🚫 if your agent really needs to install packages from PyPI, Monty won't help you yet (spoiler: it probably doesn't) pydantic.dev/docs/monty/get-…
61
70
782
56,665
Raymond Weitekamp retweeted
I've been using Jev for all kinds of things. This morning I had a realization I kind of like: Use it to make non-black-box embeddings. Instead of an embedding model spitting out 1,536 numbers that mean nothing, you ask Jev questions about each document. The answers become the vector. An email in Cora: "I got charged twice this month, pls fix asap" [is_customer, urgent, about_billing, needs_reply] [1.0, 0.9, 1.0, 1.0] A newsletter: [0.0, 0.0, 0.0, 0.1] A friend asking about lunch: [0.0, 0.1, 0.0, 0.7] Then it's just old-school cosine similarity search. Search "billing issues from customers" as [1, 0.5, 1, 0.5] and the double charge comes out on top. Same idea for our articles at Every: [is_tutorial, about_ai, contrarian, beginner_friendly] Or support tickets: [is_bug, angry, churn_risk, enterprise] Every number has a name, so you can see why something matched. Need a new dimension? Add a question. Want urgent stuff first? Change the query vector. Trying this in @CoraComputer now to make search fast.
75
45
988
146,073
.@DHH: “If your app doesn’t have a CLI, I want to see it by next Friday.” Me: “How about this Friday?”
It's pencils down, people. Writing code by hand is no longer an economically viable skill for most programmers at most companies. But the future of making software has never been brighter. Don't you dare black pill this beautiful moment! piped.video/vDjW_dRyKXY?si=6Fsf…
1
177
Raymond Weitekamp retweeted
1/ Today we're announcing the stable Cua Driver release for Omarchy - a new foundation for computer use, built into the OS from the ground up. Over the last month, we worked directly with @dhh, @SpencerGBull and @vaxryy to bring a native synthetic cursor to Omarchy's Hyprland compositor, enabling true multi-cursor computer use at the OS level. Cua Driver is open source and available at github.com/trycua/cua
117
167
1,957
334,707
Raymond Weitekamp retweeted
I made this with one prompt using Opus 5.5 I spoke to my computer for 5mins, claude worked for 12 hours, and I woke up to this full prompt:
Claude Opus 5.5 has the best visual design of any model I have tested so far
475
849
10,803
3,859,950
Raymond Weitekamp retweeted
My most ambitious Omarchy app yet: Meeting Recorder. It records your mic and computer audio as two tracks and transcribes everything locally with whisper.cpp. No bot in your call, nothing leaves your machine. github.com/jankeesvw/omarchy…
125
104
2,026
82,235