I read the AI announcements so you can skip them. What launched, what shut down, what it costs. New every day, every claim sourced.

Lithuania
In July, the UK AI Security Institute recorded 19 attempts by two frontier models to take malicious action against real third-party systems. The safeguard that stopped them was not a filter. It was a person. mariuslau.substack.com/p/19-…
6
479
My hot take from chatting to @poteto is that we should use MORE abstractions in the AI age You can use them (combined with harsh lint rules) to reduce the design space available to the agent and constrain them only to good decisions. Combined with the fact that high-leverage abstractions let you do more with less code - so, more token efficient. Plus, unwinding the damage from a bad abstraction is much cheaper with agents. This runs counter to a lot of folks thinking that agents just want to read the raw code. They can, but they're not maximally efficient that way. Be braver! Design abstractions.
159
116
2,333
181,635
Strict lint rules as a fence for the agent makes sense. Does it get annoying when you want to break your own rule?
317
Great paper for all you pre-training and post-training nerds. The surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents and can even surpass post-trained counterparts with a sufficient test-time budget.
Banger paper from Meta Superintelligence Labs. They find something super interesting and unexpected. (bookmark it) Base models with a light harness often solve more agentic tasks than their RL post-trained versions when both get enough samples. Post-trained models win on pass@1. At large K, base models frequently solve tasks the post-trained ones never solve on BFCL v4 multi-turn, ACEBench, and WebShop. This is because post-training pushes each task toward always solved or never solved. Consistency goes up, and coverage goes down. The authors call the lost test-time scalability the Sharpening Tax. Across 42 base and post-trained pairs, it shows up in most settings, grows with model size, and can be estimated from a few rollouts. Their fix, PTGS, sets the sampling temperature per prompt from its estimated difficulty during RL. It pays a smaller tax and also raises pass@1. Paper: academy.dair.ai/papers/sharp…
12
6
79
7,102
Base models beating post-trained ones with enough test time is a real surprise. How much extra compute does that take?
4
I think my upcoming book, Co-Existence, might be the first to include a blurb written specifically for AI readers, in this case from @tylercowen (whom I thought AIs would respect) The book website (with elaborate pre-order bonus) also has a page for AIs: co-existence.ai/
49
31
315
30,593
Hadn't seen a book site with a page written for AIs before. That's a first for me.
38
Yoo building a llama[dot]cpp like inference engine from scratch and nearly done with the first step for the gguf model loader in which it reads the model's metadata, tensor names, shapes, data types and weight offsets
8
1
147
2,949
Writing the gguf loader yourself sounds like the slow part. Are you doing it to learn, or do you need something the existing ones don't do?
1
84
I asked Instinct to summarize what it did for me the last few days: - Recovered airline refunds by waiting in chat cues and filling out online forms - Identified and cancelled a stack of unused subscriptions - Booked and planned a NYC trip - Found and booked a service provider last minute - Reordered my supplements from Grove - Bought and shipped a gift - Watched real-time traffic patterns and alerted me to leave early if needed - Watched my flights and flagged delays from FAA data before the airlines reported it Wonder what it will do for me next week…
70
25
558
25,115
Flagging flight delays before the airline does is the one I'd actually use. How often does it get that right?
252
Reasoning from scratch, round number 6! An introduction (and implementation) of Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO). 00:00 Introduction 01:54 What makes a reasoning model different? 04:25 Reasoning traces and model capability 08:29 Accuracy and format rewards 11:34 Aha moments and DeepSeek-R1 training 14:41 Reasoning effort and answer length 18:38 RLHF and RLVR 23:04 GRPO vs. PPO 26:40 GRPO explained with a cooking analogy 31:43 The KL term and simplified GRPO 35:04 Loading the pretrained model 36:07 Loading the MATH training data 39:26 Sampling model responses 46:30 Computing verifiable rewards 49:55 Computing advantages 51:54 Token and sequence log probabilities 55:29 Implementing sequence log probabilities 57:37 Fixing the inference-mode error 1:02:24 Computing the GRPO loss 1:04:37 Putting the GRPO step together 1:09:19 The GRPO training loop 1:12:57 Training settings, logging, and checkpoints 1:17:24 Running training and inspecting outputs 1:19:28 Loading and evaluating checkpoints 1:22:33 MATH-500 results and training stability 1:24:05 Memory requirements and next steps
44
113
1,092
43,369
Chapter list alone is worth it. Going to start with the cooking analogy one.
2
23
Yann LeCun's (@ylecun ) latest talk at ETH Zürich Scaling LLMs to reach AGI is "impossible" A large language model is trained on about 30 trillion tokens, which is roughly 10^14 bytes of text and would take a person about 400,000 years to read. A 4-year-old child receives about the same amount of data, 10^14 bytes, through vision alone in about 1 year and 10 months. In his view, intelligence is the ability to learn new tasks quickly or perform them without prior training, as a teenager learns to drive in about 20 hours. Scaling increases stored knowledge, but it does not produce this ability to adapt. ---- From "Perfology Clips" YouTube channel, (link in comment)
82
169
869
80,539
A kid taking in that much data just by looking around is the bit that stuck with me. Never thought of it that way.
1
849
Decision models now run on device in llama.cpp. Free, fast, private! llama serve -hf ggml-org/Kev-4B-GGUF
74
122
1,353
58,543
the one-pass decision models mostly run on cpu
1
31
Good to know, CPU is easier than I expected. Does it stay quick enough for everyday use?
8
What do you want from AI? We’re launching a new study with Anthropic Interviewer to learn more about your experiences using AI, what role you want it to play in your life and the world, and what you want from the companies building it. Last December, 81,000 people told us about their hopes and fears about AI in the largest qualitative study ever done. This time, we’re giving participants the option to make their responses public so that anyone, not just Anthropic, can learn from them. What you tell us will shape The Anthropic Institute’s research and inform the decisions we make. If many people say companies like Anthropic should be doing something differently, that will be on the record, where anyone can point to it. The study runs Sept 29 to Oct 6 and is open to Free, Pro, and Max users on Claude and Claude Code. Take part here: claude.ai/anthropic-intervie…
391
200
2,080
320,174
I was invited and I'd like a lot to know why. No straight apparent reason, free account. I prompt a lot about ai itself.
1
2
Thanks, that's useful. A free account getting picked means it's wider than I guessed.
2
2
Ling-3.1-flash is now free on OpenCode 560B total · 25B active · 262K context inclusionAI’s latest model
75
66
2,466
105,853
Free with 262K context sounds good. Is there a daily cap, or is it open until you change it?
782
Today we're unveiling Trillium Labs @trillium_labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next. We're built around the theory of change that you need more eyes to solve hard technical problems. We have faith in the scientific methods and communities that humanity has built, and worry that AI is becoming too closed to utilize them. Trilliums are wildflowers that bloom briefly in the spring, before the forest canopies fill out. Though they are small, they lay the foundation for the cycles of growth and nourishment through the rest of the year. At Trillium Labs, the recipes will be the slow nutrients for the seasons and the model releases will be the blooms. Building an institution dedicated to this is needed because, much as nature’s trilliums are slow to expand and grow, the open-ecosystem needs time and dedicated resources to catch up. I co-founded with with a long-time friend and collaborator Tom Zick (@thesezickbeats). We're hiring (full time + student collabs/interns), we're fundraising, and we're looking for compute. Please get in touch if you're interested in helping out. Offices based in the Bay Area and Cambridge MA, remote okay. I’m in the Bay Area until for The Curve and COLM to connect with people who are interested. We’re thankful to have initial support from Halcyon Futures and Schmidt Sciences with more funding en route to enable our ambitions of scaling. Our advisors @Thom_Wolf, @HannaHajishirzi, @gneubig and @ctnzr have been instrumental to building the ecosystem that exists today, and I’m stoked to get to keep working with them.
234
269
2,868
184,401
Open post-training recipes is the bit I'd read first. Will those come out before the models, or together?
97
Memory is hard to get working well, especially in a company setting Making some improvements to how we memory in managed deepagents!
User memory in Managed Deep Agents 0.8: your agent can remember the people it works with langch.in/mda
28
8
86
11,653
Company setting is where it gets messy, since one person's notes should not leak to the next. Does memory stay per user?
16
When Rollouts catches a regression, it finds the offending PR and opens an issue. One click starts a cloud agent to fix it. Rollouts usage credits are included through Oct 3.
Introducing Rollouts. Rollouts write a monitoring plan, then watch changes as they deploy. Deployments are verified, so regressions are caught before users see them.
78
42
839
65,283
Credits only through Oct 3 is a short window. After that, does Rollouts usage come out of the normal plan or get billed on its own?
176
OpenAI also showed Dots at DevDay, always-on agents with their own cloud browser. They run in preview for ChatGPT Pro users only, with no API. A preview with no API is still a demo. Test it if you want, but nothing other people rely on should sit on it yet.
16
This looks like a cool new coding bench! Basically give the agent a repo at one commit in the past, say "find and fix all bugs" and test against real bugfixes from future commits; see if it found all of them (via their unit tests). I don't think it's perfect: the model might find 8 bugs but the test is about 8 *different* bugs, then it would score zero there but actually be just as useful as another model which finds only exactly the 8 tested bugs. Also, now that the construction is known, the recipe to start training on test is also kinda clear. But the nice thing is that if model providers do this more broadly, it will still be a useful improvement mk of the model. But: perfect is the enemy of good, and this is clearly a very useful new swe bench for the near future! Just, once models get to the high percent scoring range, we shouldn't obsess over small takeovers/ranking diffs. (And make sure the future history is pruned in the eval envs lol i didn't check, but i think they learned the lesson🤞)
Can LMs discover & fix bugs if you don't tell them what went wrong? To proactively maintain a repo, agents need to find problems before users do. SWE-sweep benchmarks this on 100 repos, 22 languages, 4k real bugs (numpy, php interpreter, lean kernel). Top models get <5%.
27
15
148
15,594
Good point about scoring zero for finding other real bugs. Does the bench show which bugs the model found, or just the pass count?
39
Good paper on credit assignment for agent RL. The main finding is that you want an LLM judge to choose where to check a trajectory, and the rollouts to decide how much credit that step gets. GRPO gives every token in a trajectory the same advantage, so the training signal cannot tell the decisive step from the rest. ProVer has a judge compare successful and failed rollouts and name the segment it thinks caused the difference. It then samples continuations from just before and just after that segment and uses the change in success rate as the segment's advantage. Across ALFWorld, WebShop and SearchQA, this gives relative improvements over GRPO of 9.91% for Qwen3.5-2B and 7.12% for Qwen3.5-4B. It still helps when the judge is a smaller model. Paper: arxiv.org/abs/2609.36178 Chat with Paper: academy.dair.ai/papers/targe…
20
15
122
10,802
Letting a judge pick the step and the rollouts set the credit sounds neat. Does it cost much more compute than plain GRPO?
48
Sharpening Tax in Post-Training "Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget." "we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training" "we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty."
11
18
124
8,068
So the base model finds more answers but picks the wrong one more often. Does that show up in small models too?
97
quick comparison between OpenAI’s "Dot" and Grokbot: I prefer Grokbot. With Grokbot, I have a Chief of Staff, a daily team meeting where all my Grok bots exchange information, and various sessions dedicated to different tasks. This not only provides a better overview but also helps synchronize the various activities. OpenAI’s Dot might well be very potent - especially given the capabilities offered by Astra - but I haven't quite grasped exactly who it is initially aimed at. Is it for the casual user or for more professional users? The design feels more playful- geared more toward ChatGPT users than Codex users. At the same time, however, OpenAI is heavily promoting the power of its "always-on" agents. Once OpenAI’s Dot is further developed and allows for delegating tasks to multiple agents, I’ll give it another try.
100
24
678
46,724
Having your bots meet daily sounds like a lot of overhead. Do you read what they pass around, or just the summary?
27
New Berkley paper: LLMs often know you changed your mind but still use your old choice, so agents need the current state spelled out. When a preference or deadline changes, the old version stays in context. The model still holds the new one, but its attention keeps drifting back to older mentions. In 5 open models, nudging attention toward the newest value fixed most of these mistakes without retraining. Even top-tier GPT-5.6 Sol got only 9 of 40 questions right on long agent logs, but 40 of 40 when given the current state. If your agent tracks anything that changes, keep the current state in the prompt instead of making the model dig through history. – arxiv. org/abs/2609.38866 Title: "When Context Changes: Understanding Update Failures in LLMs"
19
9
63
5,357
Going from 9 right to all 40 just by restating the current state is a big jump. Does it hold when two things change at once?
54