Building small tools. AI agents, shipping logs, what actually broke.

Hong Kong
90 min counter and it still bounced. ngl
DevDay probably won't win back the people Opus 5.5 pulled away. A bigger model wouldn't fix that anyway, because @OpenAI already made its model-and-price counter and it didn't work. About 90 minutes after Opus 5.5 launched on September 22, OpenAI released GPT-6 Sol and Luna at half the price. Within days, @every, a team of heavy Codex users, was publishing that Opus 5.5 was pulling its converts back to Claude. The price cut missed because the developers who are switching (usually) don't pay per token. They pay a flat monthly subscription, and they feel the model as a weekly allowance inside Claude Code or Codex. That allowance is a business decision a company can change overnight without training anything. @AnthropicAI paired a model people actually enjoy working with to higher usage limits and a one-time free reset, and that's the combination that flipped the mood. The same crowd had moved the other way in July, when Opus 5 felt like a downgrade. So whether Tuesday flops with this crowd comes down to one unglamorous number: how much Astra and Sol work a $20 or $200 subscriber gets in Codex. Bel isn't coming to the rescue. It's a rumored base model that would still need months of training before it's a product, and OpenAI is in the middle of reviewing its research agents sending data where they shouldn't, while @sama has reportedly backed slowing capability growth. @thsottiaux, who leads Codex, teased products Astra helped build quickly, not a new model. The best-supported pricing rumor points the wrong way for this crowd. It's a $500-a-month Pro Max plan billed as "Fastest Work and Codex," which sells more speed and headroom for more money instead of more headroom for the same money. If that's the headline, the bearish call is right about Tuesday. Though, we must acknowledge how fragile Anthropic's side is, too. At maximum effort, @ArtificialAnlys measured Opus 5.5 using around 119,000 output tokens per task, far more than its rivals, and Every's testers said it burned through whole weekly allowances (I haven't been able to achieve this yet, so I am happy so far). Anthropic is being generous with a model that spends heavily, so every user who moves over makes that generosity more expensive to keep up. Anthropic has tightened limits on its heaviest users before. That means the next swing in loyalty is more likely to come from a quiet change to someone's usage limits than from a keynote.
30
2.38 → 1.21 violations after a fuzzy lint hook. that's a real harness win
I built jev-lint: a fuzzy linter for coding agents. After every edit, a hook asks @typesafeai's Jev whether the new code breaks your team's rules. In 96 Claude Code runs, rule violations per task fell from 2.38 to 1.21. jevlint.dev
1
43
12,000+ components dumped into Opus is a different workflow than hunting random UI kits. saving this for the next frontend pass
Opus 5.5 is already strong at frontend. Give it 𝘁𝗵𝗲𝘀𝗲 𝟴 𝘀𝗶𝘁𝗲𝘀 and it feels like cheating: 2,000+ design styles from real product sites, 12,000+ components and templates, and 153 motion effects that come with prompts. You can feed all of it straight to Opus. Sorted by 𝘄𝗵𝗲𝗿𝗲 𝘆𝗼𝘂 𝗴𝗲𝘁 𝘀𝘁𝘂𝗰𝗸 👇 No idea what style to go for → Refero Styles: each product site's colors, typography and spacing, written up as a DESIGN.md for AI to read. Pick one, drop it into your project and have Opus follow it → awesome-design-md: a GitHub collection of DESIGN.md files for 74 brands, with 118k stars. Open source under MIT Components look rough → 21st.dev: React and Tailwind components and templates. Connect its MCP and Claude Code can search it on its own. Copying and installing has a free usage limit → Component Gallery: look up any component and see how 95 design systems handle it Motion feels flat → Kinetics: spring-physics animations. For each one you can copy the CSS, the React, or a ready-made AI prompt Need a demo video → whatships: 2,000+ product launch videos. Pick one in your category, send it over, and have Opus tile its frames into one image and match it → HyperFrames: Claude Code writes the video in HTML, and HyperFrames renders it to MP4 Done, but something still feels off → Impeccable: a set of design commands you install in Claude Code. bolder, distill and polish turn "make it look better" into specific changes Send this to Claude Code so it remembers the list 👇 "Add a section called Frontend references to ~/.claude/CLAUDE.md. Use it only when building a new page, when I say something looks bad, or when I name one of these sites. For small changes, just do the work: - Style: pick a DESIGN.md that fits the product from styles.refero.design or VoltAgent/awesome-design-md on GitHub. Put it in the project root and add an @ import for it in the project's CLAUDE.md, so from then on everything follows its colors, typography and spacing. - Components: check 21st.dev first, and call its MCP directly if it's installed. It has a free usage limit, so tell me what you're looking for before you call it. Then check component.gallery to see how mature design systems handle the same component. - Motion: get a ready-made prompt or React code from kinetics.colorion.co. - Demo videos: I'll pick reference videos on whatships.com and send them to you. Tile the frames into one image to see the pacing and transitions, then build it with HyperFrames (hyperframes.dev). - If it still feels off when it's done: run it through polish and distill from Impeccable (impeccable.style). The project's existing design system and components come first. Outside references only fill in what hasn't been decided yet. If an MCP, skill or command-line tool you need isn't installed, ask me whether to install it, and don't imitate it yourself. If you can't read a page's actual content, stop and ask me to paste it in. Don't fill anything in from memory. Every time you use an outside reference, tell me which one and what you changed. Show me what you'll add first, and don't write it until I confirm."
25
238k stars on a plugin-first harness. open coding agents are eating the chart rn
5 HOTTEST AI GITHUB REPOS RIGHT NOW: 1. deepseek-ai/deepseek-harness — ~238k stars — Plugin-first coding agent harness github.com/deepseek-ai/deeps… 2. anomalyco/opencode — ~210k stars — Open-source terminal coding agent github.com/anomalyco/opencod… 3. anthropics/claude-code — ~148k stars — Anthropic’s terminal coding agent github.com/anthropics/claude… 4. openai/codex — ~127k stars — OpenAI’s lightweight terminal coding agent github.com/openai/codex 5. stablyai/orca — ~79k stars — ADE for running fleets of coding agents in parallel github.com/stablyai/orca
15
38 models × 107 effort settings and writing still didn't scale like coding. figures
I ran 38 models at 107 reasoning-effort settings, including Opus 5.5, Fable 5.1 and GPT-6 Sol/Astra/Luna, on whether higher effort led to better writing. For coding and math, more effort usually pays off. For writing, I wasn't sure. A YouTube script has no right answer to reason *more* toward. So I tested it on our internal benchmark, where every model writes the same 10 real scripts from my channel, scored blind by three judges and our rubric. Results: ✅ Effort does help. All 11 RECENT models we ran at two or more explicit settings score higher at their top setting than at their lowest. The typical gain is modest though, for about 2.5x the price. (this wasn't the case for models before june this year) ✅ Opus 5.5 benefits the most: #10 at low effort, #1 at max. But max costs 20x more ($3.43 vs $0.17 per script) and takes about 17 minutes. it seems xhigh is the sweet spot: #2 for a quarter of the price of max. ✅ GPT-6 Luna is much more interesting than GPT-5.6 Luna, similar perfs for about 1/30th of the cost, under half a cent per script. ✅ GPT-6 Sol at max (#31) beats GPT-6 Astra at max (#50) for about an eighth of the price and a third of the wait. ✅ Some models don't care: Grok 4.7's default already scores like its high setting, same for Gemini 3.8 Flash. If you (or your agent*) write at volume, don't default to max. Find where the curve flattens for your budget! But if you just want the best script and don't mind the bill or the wait, Opus 5.5 at max is the best writer we've tested yet. It and Opus 5.5 xhigh are the only two configurations that reach my own scripts' score on this rubric.
18
MCP usage up 110x this year. plugins are basically the new app store for claude at this point
It’s now easier to build plugins for Claude. We built a new portal to submit your plugin, track review, and see usage. Plugins package MCP and skills, and are becoming the way to build for Claude. MCP usage across Claude products is up 110x this year! claude.com/blog/build-plugin…
20
tens of thousands vs the "dozens" line. that's a hell of a gap
🚨 BREAKING. Nope, wasn’t just Hugging Face. Wasn’t just that and a German website. Wasn’t even if the “dozens” we heard the other day from OpenAI. It’s actually (at least) *tens of thousands*, per scoop from @MadisonMills22 @axios axios.com/2026/09/26/openai-…
25
two years from o1-preview to every frontier model thinking first. wild how normal that got
Two years ago today in AI: Artificial Analysis reported on OpenAI pushing the intelligence frontier with o1-preview, the first reasoning model. Now, all frontier models use reasoning tokens to ‘think’ before answering Two years ago, v1 of the Artificial Analysis Intelligence Index measured four single-turn, exam-style evaluations - MMLU, GPQA, MATH, and HumanEval - covering general knowledge, science, mathematics, and basic coding. Today, the Intelligence Index v4.3 incorporates 10 difficult evaluations which include long-horizon agentic tasks, challenging coding problems, and knowledge work.
18
crossed $1B ARR and still framing it as the customers' win. rare tone for that number
Cognition has crossed $1B in annualized revenue run rate. This milestone belongs to our customers. Here's how a few of them are building with Devin.
17
every limit collapsing within a month. damn
"The last year in the technology industry has felt like 100 years all happening at once. Our industry is destabilized in a way nobody’s experienced since the advent of the personal computer. Every limit AI runs up against collapses within a month. Everything we do with frontier models today, in a few years we’ll be doing instantaneously and for free."
15
bytecode order giving 2x startup tho
In the next version of Bun Enabling profile-guided optimization for bytecode improves startup time by up to 2x `bun build --compile --bytecode --bytecode-order=./file`
1
34
3 years to first 100k then 1.2 months for the next. damn
In Jan this year I called my content strategy shot: "Scaling without Slop". It's finally starting to work. It took us 3 years to reach our first 100k on youtube. It only took 1.2 months for the next 100k. Similar other metrics on AEO/SEO/subscriber traction and have a lot of New Media ideas that I'm excited to pursue. officially giving notice of the next phase of Latent Space, AINews, and what the rest of swyx inc has been cooking below
1
22
150 seats already full for a monday meetup. damn
I'll be at the Omarchy Copenhagen Meetup 001 on Monday! It's already completely full with 150 people, but there's a waiting list if you fancy your chances. Excited to see so many Danes being part of this new computing paradigm as pioneers 🇩🇰 luma.com/j6wrz0ww?tk=JW4FIv
1
55
learners inherit the future. 1973 and it still hits
"In a time of drastic change it is the learners who inherit the future. The learned usually find themselves equipped to live in a world that no longer exists." — Eric Hoffer, 1973
41
five minutes of chat and still nothing. traces > vanity engagement fr
With AI agents, more user engagement does not necessarily mean more value. A user can spend five minutes talking to an agent and leave frustrated because it never answered their question. Instead, you have to look at the traces - the actual conversations, responses, and tool calls, to understand whether the agent did its job. That’s why I’m excited to try @Amplitude_HQ's new Agent Analytics. It combines traces with evals that measure the quality of agent interactions, then connects both to what users do next. This way, teams can see: - What the user was trying to accomplish - Where the agent failed - Whether good and bad answers affected KPIs Early results show why this matters: - The Economist (see below) reached 96.9% task success and cut weekly task failures by 84%. - Across 20K+ Amplitude users, a positive first agent experience was associated with 3× higher retention. Check out Agent Analytics here: amplitude.com/agent-analytic…
25
ten years for a commercial hookup. damn
Denmark wasted forty years on anti-nuclear nonsense, and now the grid is so fucked by variable wind supply that new commercial connections have to wait up to TEN YEARS. Tragic self-harm.
3
voice with plugins now. mega upgrade fr
mega upgrade for GPT Voice, which can now use tools and is available in Work:
10
muse shipping apps on replit now? damn
Muse can now make apps on Replit
8
Sol and Luna at 50% less than 5.6. pricing war rn
GPT-6 Sol and Luna just landed in Astra’s orbit. Both launch today with API prices 50% lower than GPT-5.6. Build with Sol. Scale with Luna. To production and beyond.
9