AI Researcher | Exploring Agentic Workflows | Follow for daily AI insights

US, Florida
Boris Cherny: "Overnight sub-agents do deeper work" 10-step Fable 5 setup you can copy this week: 1. Write CLAUDE.md - stack - commands - code style - forbidden files - review rules 2. Add PROJECT_MEMORY.md - verified facts - failed attempts - last session - next run 3. Create 1 Skill per repeated workflow - CI triage - PR review - design QA - deploy check 4. Add eval cases Put them in: eval/<workflow>.jsonl 5. Split maker and verifier Maker writes the chang Verifier runs the app, tests, screenshots, logs 6. Use worktrees for parallel runs No shared checkout No file collisions No mystery edits 7. Send work by price Fable 5 plans across days Sonnet 4.6 does bulk edits Haiku 4.5 grades Opus 4.8 handles fallback cases 8. Put UI work behind screenshots If the task is visual, text logs aren't enough 9. Move long jobs to Routines CI failed? Run triage PR opened? Run review 7am? Send digest 10. End every run by writing the lesson back A fix that stays in chat dies there The rule: > Builder makes the change > Verifier checks the real artifact > Memory keeps the receipt > You read the diff Skip the verifier and you don't have an agent system You have a very confident intern with shell access
43
84
755
142,941
Alephic co-founder Noah Brier: "There's entirely too much focus on [AI's] ability to write" He shows how Claude Code reads his whole Obsidian vault, pulls the notes he needs and catches him up on his own research Watch it, then copy the CLAUDE.md manual and the six slash commands from the article below
3
4
501
Morning Brew co-founder Alex Lieberman: "I have basically trained it to only use my words" He shows how his Claude content machine writes only from his interview answers and learns from his edits Watch it, then copy the writing-guide and feedback-loop prompts from the article below
6
9
759
TypeSafe CEO Diogo Almeida: "I want to automate the easy work before the hard work" He explains why routine branching belongs to a typed decision model, not a frontier LLM Watch it, then copy the Jev -> Sol -> Astra ladder and the approval gate from the article below Full video: piped.video/cFx9Z3ZXca0
5
745
Serval CEO Jake Stauch: "People remain the biggest moat you can have" He explains why AI companies are hiring more than ever, even though AI was supposed to take the jobs Watch it, then steal the 4-point Automation Engineer framework from the article below
3
5
684
OpenAI Head of Core Products Tibo Sottiaux: "The agent never stops, it occasionally sleeps" He explains how a Dot learns from your feedback, and how a second Guardian agent stops risky actions Watch it, then copy the Dot job prompt and permission rules from the article below
OpenAI just opened a new era for the AI industry Dots, Sol 6.1, a decision brain and more: 20+ life-changing products that are reshaping AI engineering Use them correctly, and you’ll get a second brain In this article, I show you how
Article

OpenAI New Era - Set Up and Use: Dot Agents, Sol 6.1 and More in 10 Steps (From Scratch)

99% of people didn't realize that OpenAI has made yet another breakthrough in AI engineering - Completely new (hello, Grokbot), life-changing launches that I'll discuss in this article: what they

6
1
12
746
Remotion founder Jonny Burger: "The time is the argument and your job is to return an image" He explains why a video has to be a pure function of time, and why CSS animations flicker in render Watch it, then copy the seek(t) renderer and CLAUDE.md render contract from the article below
6
14
792
Remotion founder Jonny Burger: "You can use actual programming, a real programming language, to define the content of your video" In this interview, he shows how Remotion turns React, APIs, and agent skills into repeatable video workflows The guide below turns the same idea into a 5-file Opus 5.5 studio, with exact prompts and $26.77 spent before the final render Watch the interview, then copy the 5-file setup below
4
1
15
792
Sam Witteveen: "They're often just really simple classification problems" In 16 minutes, he explains Choice, Score, and Noul, then runs a support-ticket router Watch it, then build your first Jev request with the article below
1
1
13
799
AWS Developer Advocate Mike Chambers: "Put those gated decision points inside of an agent flow" He shows how Jev can block risky tool calls before they run Watch it, then copy the Claude Code safety gate and stop hook from the article below
3
1
9
790
Mastra co-founder and CTO Abhi Aiyer: "There's a lot of stuff in workflows that have to make a decision. Usually we tell people to make an agent step to then decide what the flow should be of the workflow going forward." In this 30-minute live build, he tests Jev on routing, tool search, approvals, guardrails, and evals, and shows where it isn't worth using Watch the demo, then steal the agent-loop map from the article below
4
12
767
Boris Cherny: "Whenever we talk about verification, people are thinking like unit tests, lint, or type check. But actually, when we talk about verification for agents, it's something slightly different. It's like, can the agent run the thing?" In 18 minutes, he and Cat Wu show how Claude tests its own work, catches edge cases, and turns failures into reusable skills Watch it, then use the article below to run drafts on low effort and verification on high
What is effort really? When do you change it it and why not just use max effort for everything? I dove deep into this problem, looking into evals and doing my own tests and I was quite surprised by the results.
Article

Using Claude Code: Spending your effort

One of the best parts of our newest Claude models is how they respond to effort without breaking the prompt cache in Claude Code, but I’ve received a lot of questions on this from users. What is

4
9
801
Google ML Developer Expert Sam Witteveen: "The challenge with most of those is that each one needs to be trained or fine-tuned on your specific use case" In 21 minutes, he compares 7 open Jev-style models and shows where their training holds up or breaks Watch the comparison, then use the article below to fine-tune Qwen3.5 4B on about 38K examples for a $17 training job
4
14
627
SE Ranking’s Guifré Ballester: “You give it an objective, a goal, and it figures out the path. It plans the steps, executes them, reads and writes files, calls APIs, and keeps going until the job is done.” In this 69-minute workshop, his team connects Claude Code to live SEO data and builds a content brief, an AI-search visibility check, and a backlink script on screen The article below shows how to run the same loop every Monday and keep the agent focused on pages that bring signups
4
19
734
TypeSafe founder Diogo Almeida: "I want to automate the easy work before the hard work" In this 2-hour interview, he explains how Jev uses choices, scores, and yes/no probabilities to route work inside agent loops The guide below has 3 builds: a skill router, a codebase linter, and a retrieval reranker Watch the interview, then copy the builds below
6
15
830
35 AI researchers just dropped a 64-page paper on Graph Engineering. It's a map for turning one messy agent loop into explicit tasks, specialist teams, checkpoints, recovery paths, and verification Drop the PDF and the article below into Claude Code or Codex, then ask it to turn your longest workflow into a graph you can pause, inspect, and resume. Full paper in the comments 👇
4
10
704
ChatGPT co-creator Diogo Almeida: "AI: too good to be true, too bad to be useful." This 36-minute lecture is basically the guy who helped build ChatGPT explaining why chat models make sketchy employees. He spent 4.5 years at OpenAI working on InstructGPT, GPT-4, ChatGPT, and RLHF, and now he's building Jev for the tiny decisions those models keep overcomplicating. The article below shows the working setup: Jev handles the quick calls, Kimi K3 gets the uncertain cases, and code keeps permission to act. Watch the lecture, then steal the 7-day build plan below.
5
10
780
Grok 4.7 just landed at the same $2/$6 API price and speed as 4.6 SpaceXAI’s published evals show the biggest gains in agent work: • CursorBench: 40% -> 46% • DeepSWE: 65% -> 71% • EEBench: 53% -> 64% • Terminal-Bench: 20% -> 38% • Harvey Legal: 16% -> 20% It beats GPT-5.6 Sol on 5 of 7 listed benchmarks, while Fable 5.1 still leads 4 of 7 Artificial Analysis scores Grok 4.7 at 46, up just 2 points from 4.6 Verdict: a strong efficiency release with an incremental intelligence gain
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
7
1
17
750
Codevolution creator Vishwas Gopinath: "The simplest way to think about it is as a smart if statement" In 22 minutes, he explains Choice, Score, and Noul, then builds the same decision flow in the Playground and TypeScript Use code for the obvious calls and Jev for the messy ones The article below gives you 6 workflows to test that split yourself
9
14
807
Towards AI co-founder Louis-François Bouchard: "The research needs flexibility, but the writing needs constraint. And so, the decision is to split into two systems" In this 2-hour workshop, his team builds a research agent that searches, inspects sources, finds gaps, and saves the evidence for a separate writer Sending all 50 sources straight to the writer buries useful papers beside weak matches The article below shows how to score all 50 sources before the agent writes a word
5
1
15
1,095