NVIDIA and Arcee just explained why companies moved to open models 43 minutes from the AI Engineer World's Fair stage They wanted a model they could run themselves and always have access to Jev signups are paused now, and open copies like Kev already run on a laptop GPU Watch it, then read the full guide on running open Jev models on your own GPU below
8
17
1,620
AWS just showed how Jev stops an agent from calling a tool with a made-up argument 11 minutes from Mike Chambers, senior developer advocate at AWS One request asks Jev 4 yes/no questions about the tool call before it runs Jev scored the assumed Seattle 0.04, so the agent asked which city the user meant Watch it, then read the full guide on writing Jev questions below
4
18
2,071
Real Python just swapped a yes/no if-statement for one Jev question 17 minutes of Python you can type along with A plain yes scored 0.34 until they reworded the question, then it scored 0.98 Any score between 0.2 and 0.8 makes the script ask again Watch it, then read the full guide on your first Jev request below
4
18
2,468
Researchers just caught AI agents inventing answers when their tools quietly break They gave an agent over a thousand requests, like a bank balance or a patient's lab result, and rigged every tool to send back nothing usable When the tool openly returned an error, the agent reported it honestly But when the tool said everything was fine and sent back a blanked-out value, the agent answered as if it had the number, or made up a reason it couldn't share it, almost half the time What makes it worse is that every agent framework prompt they tested had this problem, and none of the nine they checked tells the model what to do when a tool fails Personally, I'm adding one sentence to every agent I run, it has to mark each lookup as OK or FAILED before answering, and in the paper that alone cut dishonest answers from 14% to under 1%
8
3
16
1,741
Duke researchers just showed you can break an AI search agent without feeding it a single false fact They slipped true pages into its search results, each one answering a slightly different version of the question Asked which film won Best Picture at the 89th Oscars, the agent read a true page saying La La Land was first named the winner, and answered La La Land instead of Moonlight And the later that page showed up in the search, and the more it looked like a finished answer, the more often the agent took the bait The worst part is that fact-checking can't catch this, every planted page is true, and telling the agent to double-check every condition in the question only partly helped Personally, this is what worries me, we're handing our research to agents that can't tell a true answer from the right one
5
1
16
1,710
Exa CEO, Will Bryk: "AI systems really want perfect search. They don't want SEO. They don't want ads" He expects AI searches to pass human searches this year and hit 1,000x within a few years 17 minutes on why search built for people fails the agents that will soon do most of the searching Every weak page an agent reads still eats its token budget So I wrote a guide on screening 50 sources with Jev before your agent writes a word It also shows how to catch the good papers the filter throws out Watch it, then grab the starter and prompts below
1
12
1,970
Mastra CTO, Abhi Aiyer: "If I had 200 tools and I asked about refunds, the classifier would say this dude's asking about refunds, that's the refund tool domain" 16 minutes of live code wiring Jev into Mastra: workflow branches, guardrails, tool pre-selection, classifier judges His own tests put Jev 3rd out of the tool search options he tried So before a Jev filter decides what your agent reads, count the useful sources it threw away Watch it, then read the guide below to put a Jev screen in front of 50 research sources, with a review queue and a count of the useful papers it archived
6
1
14
1,912
Teams still pay a heavy LLM to make calls a panel of humans would answer in 5 seconds LangChain sat down with the team behind Jev for 46 minutes on where those calls belong In the live demo an agent tries to delete a customer, Jev marks the call risky, and the tool never runs A second middleware sends the easy prompt to a fast model and keeps the big one for the hard task Their rule for scoring: describe every level well enough that a 1.5 means the same thing on every document you score And the cutoff stays yours to prove, because the confidence you get back is a statistic over the probability map, not a silver bullet My guide turns that into a research filter: 50 papers, a 0 to 3 relevance rubric, a review queue, and a check on the useful papers it archived Watch the session, then grab the starter and the prompts below
3
1
11
2,274
LangChain put Jev inside their own agent loop and showed the three jobs they hand it In the demo an LLM took about 5 seconds to say whether a text contains PII, and Jev answered immediately with 98% So their coding agents can ask Jev how hard a task is and save the powerful model for the hard ones Their auto mode middleware asks Jev whether a tool call is risky and blocks it at runtime, a check she had switched off because it used to be too slow And the same model grades an agent's answer against a rubric in online evals: correct, matches the reference, grounded, cites a source My guide points that setup at research sources, with a 0 to 3 relevance rubric over 50 papers and a review queue for the low-confidence calls It also checks which useful papers the filter quietly archived Watch the session, then grab the starter and the prompts below
3
15
2,480
Your research agent can miss the best paper before it writes a word Joe Maddalone explains where he'd put Jev in his workflow: score incoming ideas before passing them to an LLM for development Before trusting that approach for research, check how many useful papers the filter rejects Watch this 62-second clip, then save the guide below for the Python workflow to screen 50 papers and send uncertain cases to manual review
3
18
2,634
Steve Sewell, builder of Agent-Native: "Jev only makes decisions." In this 10-minute breakdown, he explains how selecting tools and skills for each prompt keeps irrelevant instructions out of your agent's context. Watch it, then use the guide below to build a Jev router that sends uncertain decisions for review.
Jev is the "Internet" moment for the AI industry It tells your agents and LLMs what to do next, in milliseconds and at almost zero cost If you set it up correctly, you will have the AI engineer’s stack for 2028 In this article, I show you how
Article

Jev Engineering: Full 10-Step Roadmap to Set Up and Use a New Brain for AI (from scratch)

The Jevons Paradox (he's on picture) is a rule stating that an increase in the efficiency of a resource's use does not reduce, but rather increases, its overall consumption - That's the global

3
13
2,447
Shopify CEO Tobi Lütke: “We have always created superintelligence around us” He means cities and societies: people combining abilities no individual has alone For AI agents, compare what each agent can do alone with what the group can do together Pascio's article argues we should watch agent traffic and capabilities that only appear across a whole pipeline Watch the 56-second clip, then read the article below for its checklist of what to track between agents
3
1
17
2,019
OpenClaw creator Peter Steinberger: "You want to extend the loop. So any input can actually be verified" 19 minutes on how he runs his open source repos with agents that check the work before he does A new issue gets checked against his vision.md, then a second agent reviews and repairs the PR, so it's ready to merge when he looks His agents "need much less babysitting because you just give them more tools to do the work" Watch it, then copy the gate and stop rules from the guide below
5
1
23
2,429
OpenClaw creator and OpenAI engineer, Peter Steinberger: "The agent runs the inner execution loop. I set the direction and I make decisions in the outer loop" When someone files an issue on his open source project, a manager agent checks it against the project's goals, a worker writes the change and runs the tests, and a second agent reviews it He skips the intermediate messages, reviews the PR once, approves, and the change lands after the checks pass 6 minutes on how he went from juggling 10 terminals to managing a manager agent Watch it, then copy the loop prompt from the guide below: success gate, blocked actions, stop rules, run log
5
20
2,285
Replit CEO, Amjad Masad: "We closed the loop and we have an agent internally at Replit that is constantly evolving Replit Agent" Every night that agent reads the traces from Replit users and opens a pull request with prompt changes for the errors and sentiment issues it finds Each PR ships as an A/B test, and the result decides whether Replit releases the change or edits it 46 minutes with SaaStr's Jason Lemkin, whose agent 10K emailed 331 investors on its own Watch it, then copy the general loop prompt from the guide below
4
11
1,792
GPT-6 Astra built a playable map for a Call of Duty clone from a PDF of an estate. Then launched a match and flew a helicopter around it to test its own work. VibeCode CEO Riley Brown ran the build in Codex and watched it catch the house coming out too white. "It's basically reprompting itself in this loop" That loop is the reason to hand Astra the jobs that keep coming back to you half-finished. Watch it, then read the guide below before you move your builds to Astra
7
3
17
2,526
LangChain engineer: "If you're using 100,000 tokens, even if it's cached, it's still 100,000 tokens" 50 minutes on why one small agent task keeps re-sending the same context in every request. Manus calls about 50 tools per run, and every one of those results rides along in the next call. Caching lowers the rate on those tokens and never lowers the count. Watch it, then read the guide below on pricing one finished task.
2
8
2,272
Most staking is a lockup with extra steps. Cosmos Hub makes you wait 21 days to unbond. Ethereum's validator exit queue has run into weeks when a lot of people head for the door at once. Sui works on 24-hour epochs. Stake now, rewards start after the next epoch. Unstake, and the SUI is liquid after the next one. Roughly a day in both directions. That's the shortest round trip I've seen on a major PoS chain, and it changes the math on small positions. Minimum stake is 1 SUI. Network APR is around 1.47% right now and floats with on-chain conditions. Sui staking is non-custodial. Your keys stay in your wallet, the delegation happens on-chain, and nobody can move your SUI but you. So the only real decision is which validator gets your stake. One that goes offline or gets slashed eats your rewards. @HashKeyCloud runs validators on 40+ chains under HashKey Group, licensed in Hong Kong. 5+ years live, 99.9% uptime, 0 slashing. Their Sui commission is 8%, against a market average of 9%, so you keep more of the rewards than you would with a typical validator. If you want to try it: 1. Open a Sui wallet (Slush or any mainstream one) 2. Go to Earn, then Staking, then pick a validator 3. Search for HashKey Cloud 4. Enter the amount and confirm Rewards start after the next epoch, about a day later. Total stake, live APY and delegator count are public here: suivision.xyz/validator/0x24…
🔷 Stake SUI with only 8% validator commission. With HashKey Cloud, you can access competitive SUI staking fees through reliable validator infrastructure backed by years of experience across leading blockchain networks. Built on secure, non-custodial staking infrastructure, HashKey Cloud provides a simple and reliable way to put your SUI to work while keeping control of your assets. Looking to earn staking rewards on your SUI? 👇 Start staking with HashKey Cloud today. suivision.xyz/validator/0x24… #HashKeyCloud #Sui #SUIStaking #Staking
11
11
1,807
I noticed the karpathy account in EvoMap AutoResearch's Contributors list. What caught my attention in EvoMap's release was how test failures shaped the agent's next attempt.
3
9
1,670
Learn how an AI tutor keeps track of what a student needs in this 63-minute talk from Towards AI. With caching enabled, full chat history beat compaction on recall and cost in their cloud tests. Document retrieval helped with smaller context windows. Watch it, then use the guide below to separate your Grok Bot's lasting instructions from facts that need updating.
2
7
1,734
Claude Opus 5 output costs $25 per million tokens on the official API The same model through DIT - $10 I ran a controlled test before I believed that One endpoint, one prompt, temperature 0, max_tokens 256, and the only variable was the model name: > GPT 5.4 came back with {"total_cents":855,"total_units":5} in 3.58s > Claude Sonnet 4.6 came back with {"total_cents":855,"total_units":5} in 3.95s Both HTTP 200, both right on the invoice math, identical JSON down to the key order Switching cost me two edits: base URL to api.dit.ai/v1 and a new key, and the OpenAI SDK stayed exactly where it was Grab a key and price your heaviest model against the official rate: Dit.ai
Same GPT and Claude models, one API key, and a lower bill without rebuilding your stack. Dit.ai puts 50+ text, image, and video models behind an OpenAI-compatible API, with rates typically 30–70% below official pricing.
5
12
1,738
The catalog includes GPT, Claude, Gemini, DeepSeek, Llama, Grok, plus image and video models. My OpenAI-compatible client stayed in place. I changed only the model ID, and Claude returned the same exact JSON. Here’s the Claude run.
2
2
639
Switching took two edits: change the base URL and replace the API key. I used that setup to send the same invoice prompt to GPT 5.4 and Claude Sonnet 4.6. Both returned the exact JSON requested. Here’s the GPT call.
1
737
Each request is matched across 160+ competing providers. Dit.ai picks the lowest-priced qualified route in real time, so the model stays the same while the route and price can change. That competition drives the typical 30–70% savings.
1
2
909
Same GPT and Claude models, one API key, and a lower bill without rebuilding your stack. Dit.ai puts 50+ text, image, and video models behind an OpenAI-compatible API, with rates typically 30–70% below official pricing.
11
6
37
28,107
HumanLayer founder, Dex Horthy: "You now have two sources of truth and it stops being useful" 89 minutes on context engineering from the engineer who wrote 12 factor agents The case he keeps seeing is a spec file and the code it describes, where editing one leaves the other quietly wrong So he deletes the research doc after every task and regenerates it, because he would rather spend tokens than trust a doc that no longer matches the code Grok Bot gives you 7 places to put a rule, a fact, a file, or a workflow, and the same drift starts the moment one item lives in two of them Watch it, then read the guide below on which home each item belongs in, plus the prompt that sorts your current setup
7
24
2,150
Make a fresh agent session pick up the work where the last one stopped Lance Martin explains the handoff in this 62-minute Chroma interview At 19:46, he describes a loop built around a task list, Git and progress.txt One agent completes a task, saves its changes and updates the progress record The next agent reads those files and Git history before taking the next task For Grok Bots, the article below includes a handoff prompt with exact file versions, passed checks, blockers and one next owner Watch the explanation, then copy the handoff template from the guide below
5
1
11
1,884
Hugging Face explains how to give AI agents long-term memory in 29 minutes. Alejandro AO shows how to extract useful facts from conversations and retrieve them when an agent needs them. The memory setup can use local models, too. Watch it, then read the guide below on building memory that improves between runs.
2
11
2,195
Harrison Chase, co-founder of LangChain, on what happened when he moved his own email agent to a new platform: "it didn't have all of my memories... even though it had the same starter prompt and the same tools... I still haven't fully switched over because it kind of sucks now compared to what it was before." Same prompt, same tools. Worse agent. In this 40-minute conversation with Sequoia he gets to the loop that rebuilds what was lost: an agent reflecting on traces from an earlier session, then updating what it keeps. Watch it, then read my guide below to build that loop without a managed service. It walks through the append-only event journal, the four kinds of memory, conflict fields like valid_from and supersedes, a 5 to 12 record retrieval pack, and the 50-case exam a candidate memory has to pass before you promote it.
1
9
2,154
OpenAI just spent 56 minutes on agent memory patterns Emre Okcular, solutions architect at OpenAI, builds it live in the Agents SDK, up to a cross session summary injected into the system prompt so a fresh agent opens by naming your machine, your OS version, and the fix you already tried then someone in the Q&A asks how you know the memory is actually helping his answer is run your normal evals with memory on and off, then build memory specific evals on a golden set of about 50 examples the session hands you the patterns and leaves you to write the exam yourself watch it, then read the article below for that exam
3
8
2,019
Matt Pocock's 96-minute AI coding workshop shows how a rough feature request becomes a reviewed, tested implementation. He turns it into a PRD, breaks the work into dependency-aware vertical slices, sends each issue to agents in isolated worktrees, then checks the commit, tests, type checks, and UI by hand. My guide below covers the full SaaS path in 5 gates: evidence, specification, build, release, and revenue. It includes tenant isolation, billing webhooks, deployment, rollback, and paid design partners. Watch the workshop, then save my guide below and use it to take one customer problem to first payment.
2
10
1,841
Member of Technical Staff at Anthropic, taught a Code with Claude workshop on agents that remember. The 29-minute session opens on the default most people ship: sessions run isolated, so whatever the agent works out in one run is gone by the next. Then he builds the fix on stage. A memory store mounts onto a session as a resource the agent can read and write across runs. A prompt parameter steers what gets written, access can be set read-only, and every memory file is versioned and editable by hand. The back half is consolidation. A dream is an async job that reads one input store plus up to 100 past session transcripts, fact-checks them, backfills dates and identifiers, and merges duplicates. An orchestrator spawns one sub-agent per session. The input store is cloned rather than written over, so the console gives you a diff to review before any of it reaches a future run. Watch the session for the shape of that loop. Then read the guide below to build the same thing without a managed service
7
1
20
2,247
Rajiv Shah's 34-minute session shows how a coding agent can work through a repo and prove it finished the job. The working loop is the section to watch twice: > 23:14 - plan before execution > 23:48 - verify each hypothesis and feed the result into the next attempt > 24:33 - use tests as the completion check > 26:21 - run code inside a sandbox > 26:34 - require approval for risky actions > 28:54 - put an orchestrator over workers and validation Watch the session, then save the quoted guide below. It gives you 7 files to build the workflow
5
1
7
2,108
OpenAI engineer Ryan Lopopolo has a rule for coding agents: When the model gets stuck, split the task into smaller building blocks. Turn each failure into context, tooling, documentation, or a test the next run can use. In this 78-minute interview, he shows how docs, tests, lints, and review agents capture quality requirements. Failed builds feed missing context back into the repo. Watch it, then save the guide below for 7 copyable templates
7
15
1,877
This 30-minute AI Engineer session by Nearform tech lead Alfonso Graziano shows how a coding agent can improve another agent through measured, reversible experiments. He starts with a golden dataset and scorer, gives the coding agent a job file with the objective, repo, metrics, constraints, and relevant files, then tests one hypothesis per branch. 03:35 - define expected outputs and tool calls 13:31 - establish a baseline, test a hypothesis, keep the gain or roll back 15:29 - preserve each run in a branch, report.md, and shared memory 19:44 - turn user traces into failure clusters 26:09 - add every confirmed failure to regression tests 28:30 - connect specs, quality gates, context, and observability In one production case, the eval score moved from 67% to 86% in about 10 iterations. The agent found edge cases, improved the system prompt and tool descriptions, and fixed tool logic. Watch the session, then use the guide below to build the task contract, project map, typed tools, durable handoff, evaluator rubric, permission policy, and controller.
4
1
13
1,915
Greg Isenberg: “You earn the software by doing the work first.” His 26-minute breakdown maps the next 30 days: choose a paid workflow, run it manually with AI and human approval, build an eval set, sell 2 pilots, then turn the repeated work into software. Watch it, then save the guide below to build the first paid version
5
13
2,140
Four AI workers can copy one bad source into an entire newsletter Hunter Gon’s desk makes every handoff reviewable by requiring a named file: > Researcher writes source-backed story cards > The human approves the story IDs > Writer drafts only from those IDs > Copy Desk marks each claim PASS, FIX, or REMOVE The Desk Editor waits for a human before the send Save the guide for the role prompts, folder tree, approval gates, and 7-day rehearsal
6
10
1,886
X raised reply-routing threshold from 100K to 120K followers. It uses the parent and root authors. The replier’s follower count is ignored. Either >120K: Grok reply ranking Both ≤120K: spam detection A ranking score cannot guarantee top position.
3
9
1,785
x402 creator Erik Reppel: "The thing that agents can do if you give them a wallet is pay for stuff" He says x402 already handles those payments tens of thousands of times a day His example gives an agent $5 a day and a $3 purchase cap, with the wallet forcing human approval above either limit Put those limits in the wallet code, where the model can't reinterpret them Watch the interview, then read Coinbase's AiFi article below
4
10
1,939
Matt’s Grok Bot roster has 15 bots 7 are still experiments, including one called New Bot that has never done a thing Before adding bot 16, use his counting test: > "What outcome would visibly stop if this bot disappeared?" Give one bot one recurring job and one output, then run it solo until it works Read the article, it includes 21 hacks plus the commands, role instructions, approval rules, and handoff files
6
1
19
2,256
Nikita Bier: “When you actually have something important to post, it ends up getting buried” He was explaining why low-value “GM” replies can hurt the next post with the same readers X predicts each reader’s actions and combines them with different weights before ranking the post for that reader Give the reader something specific to answer or send to someone else 73-minute interview is full of insights about X’s recommendation system Watch it, then save the article below to understand scoring formula
6
1,916
Anthropic engineer: "You can pick and choose whatever primitives you need ... and ditch the rest" His 27-minute workshop shows the payoff: 1 outcome definition starts the job, then Claude opens files, calls tools, and launches 4 subagents with separate context windows Nick Spisak's X Article gives you the build path: agent, environment, session, events, per-tool approvals, and the exact Claude Code onboarding command Watch the workshop, then copy the setup from the X Article below
9
3
18
2,357
X’s published defaults make a 1% copy-link chance worth 4× a 10% like chance 1% × 20 = 0.20 10% × 0.5 = 0.05 X calculates these probabilities separately for every viewer Give one specific reader something useful enough to send into a group chat 19-minute vid shows how X retrieves, filters and scores posts Read guide below, it covers weights, 48-hour cutoff, author-diversity rules and a practical pre-publish checklist
5
17
2,148
Replit President Michele Catasta on coding agents: "I think the shortest possible summary is autonomy" 24 minutes of Replit's President explaining how a coding agent does useful work for 10–15 minutes at a time Users can stop or redirect the agent in chat during a live run Worth watching before you give any agent 15 minutes of autonomy Watch the full interview, then read the 30-day Grok Bot fleet guide below
6
19
2,389

ALT Jimmy Fallon What GIF by The Tonight Show Starring Jimmy Fallon

1
2
132
TxFlow just opened $1,000,000 USDC trading campaign For 10 days, traders compete for a fresh $100K pool and up to $100 USDC per day Code TXBATCHER also gives you 5% off trading fees Your highest daily perp taker-volume tier sets the reward: $200K → $5 | $500K → $15 | $1M → $35 | $1.5M → $60 | $2.5M → $100 The campaign runs from Aug 21, 00:00 to Aug 30, 23:59 UTC Volume resets at 00:00 UTC, so each day gives you another shot at the pool Click Join before trading Only manual taker volume counts, positions must stay open for at least 1 minute, and Fee Credit volume is excluded Rewards don't stack, and the daily pool is paid from the highest-volume traders downward until fully allocated Rewards are sent within 5 business days after the campaign ends Open the campaign with TXBATCHER applied for 5% off trading fees: app.txflow.com/campaign/dail…
11
2
19
1,742
Technical lead at JetBrains: “Naive prompting gives you hope, domain modeling gives you a contract that you can trust” Koog grew out of JetBrains’ work scaling agents for millions of users In 50 minutes, he builds a Java banking agent with typed steps, limited tools, code-owned routing, verify-and-repair loops, and per-node crash recovery He also reports a 6–8% benchmark gain from history compression that preserves named facts Harry’s later article explains 5 graph patterns and a 7-step GitHub issue-to-PR build with tests, retries, human approval, and an exit Watch the JavaOne session, then check the article below
9
18
1,942
Replying to @vikktorrrre

ALT Happy Dog GIF

1
1
70
Trade on BloFin and qualify for up to $1,000 in Futures Bonus Only the first 100 eligible traders are included in the BloFin × Alpha Batcher campaign Campaign window: Aug 18, 10:00 to Aug 31, 10:00 UTC - net deposit tiers from 100 to 2,000 USDT: $20 to $300 Futures Bonus - taker volume tiers from 100K to 15M USDT: $20 to $1,000 Futures Bonus BloFin combines spot, perpetual futures, copy trading, and trading bots on one platform, with Merkle Proof of Reserves and stated 1:1 backing of user assets The deposit reward requires holding the funds for at least 3 days and completing 50K+ in taker volume Only taker positions held longer than 1 minute count, withdrawals during the campaign disqualify participation, and the bonus expires after 7 days and cannot be withdrawn Sign up through my partner link and complete campaign requirements to qualify: partner.blofin.com/d/alphaba…
8
13
1,743
Alex Rinke, Celonis co-founder: "If you just use an AI agent here or there, it's not really going to move the needle in terms of business outcomes" In 17 minutes, he explains why AI pilots stall and what has to exist before an agent can change a business outcome Luke Pierce turns Rinke's process requirements into a 6-step build order for $2M-$10M companies Watch Rinke for the 3 requirements, then read the guide below to map workflows and place agents after the data model is ready
8
1
14
2,271
He updated his Grok Bot machine on 14 August and only one folder came back /workspace survived, $ HOME and apt packages and /usr were gone, Tailscale and ssh needed reinstalling Every CLI your routines call installs into one of those wiped locations by default, so they start failing the week after an update and nothing tells you why He also found sudo -n id returning uid=0, since the box ships with NOPASSWD ALL Two marker files and one update settle it on your own box, and the guide names the four folders to move into
3
17
1,787
After 9 rebuilds of BabyAGI, Yohei Nakajima now builds the agent around its event log 17 minutes on the runtime he built, where replay, rollback, and fork come free 2:07 - what changes when log becomes the source of truth 3:39 - policy that makes a prompt edit wait for a human 7:10 - behaviors that watch graph rather than call each other 9:05 - run that died at question 350 and picked up at 353 11:43 - static and sandbox gates a self-edit has to clear A plain checkpointer gets you most of this, once you decide where each step ends Watch it, then read the guide below on when a graph is worth building and where to cut the boxes
1
6
1,728
Replying to @kimmonismus

ALT Jimmy Fallon Banger GIF by The Tonight Show Starring Jimmy Fallon

95
Her sidebar looks like a team roster where every name is a bot Forecast reports all 17 Next Steps updated in Salesforce, slides has the deck rebuilt from a Granola transcript, and Chief of Staff opens with inbox is quiet, nothing needs you Buried in her prospecting prompt is the discipline most people will skip, check sent mail from the last 90 days and CRM activity, then mark anyone already touched as Skip Draft It also tells the agent to write no verifiable recent posts found rather than inventing anything, and to show 2 or 3 drafts for approval before generating the batch Onboarding runs through a Teach a task button, you demo the workflow on your own screen once and the recording becomes a skill every agent can call The article carries all 5 prompts in full, the video is a 32 minute walkthrough of how skills and shared memory work underneath
2
10
2,387
Claude is about to start watermarking every text it writes It is EU AI Act compliance, and the other major labs signed the same Code of Practice, so their watermarks are coming too Finding the mark only proves Claude was involved somewhere, not whether it wrote the text or just tidied yours Light proofreading may leave nothing to detect since the words stayed yours, a translation is marked in full because Claude picked every one, and code barely registers Nothing in the key points at you, no user, no organization, no chat Anthropic names the method as a version of DeepMind's SynthID-Text, says a detection API is coming, and lists the tells keyless detectors use, including AI fondness for the word quietly
We’ve written an FAQ to answer some of the questions we've received about watermarking. In summary: • We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking; • Our watermarking method doesn’t have any practical impact on the quality or content of Claude’s outputs; • The difference between watermarked and un-watermarked text will not be distinguishable to readers; • Nothing is added to the text and there are no hidden characters; • Watermarking doesn’t require extra tokens, and will not be more expensive; • Watermarks can’t be traced to a specific person, organization, or chat. Read more: anthropic.com/news/claude-te…
4
1
12
2,122
Andrej Karpathy: "Delete everything, keep attention" He spent 68 minutes explaining the transformer as a graph Attention is message passing between nodes on a directed graph, where each node holds a private vector and emits a query for what it wants, a key for what it has, and a value for what it will send He writes one round of that communication in plain Python on screen On the 2017 paper that dropped the RNN scaffolding, his line is "delete everything, keep attention" And on shapes, RNNs are a long thin compute graph that gradients struggle through, transformers are a shallow wide graph where supervision reaches the input in few hops The guide below is a graph of the other kind, the route an agent takes, and it comes with a copy-paste Claude Code prompt and the config to run it
3
1
13
2,687
DoorDash's 130,000 automated tasks a month are dominated by one workflow Their own chart puts PR review and merge about three times higher than alerting, Jira triage or scheduled monitoring 41 minutes from LinkedIn's platform team lands on the same wall, most agent failures come from missing or stale facts, not from a model that reasons badly So both teams built the same thing, a gateway handing agents code search, logs, ownership and PR history under scoped permissions They split on the rest, DoorDash wrote Flux in-house, LinkedIn closes with buy or extend and points at the Copilot they wrapped in their own MCP servers
7
18
2,974
I get paid if you sign up, so here's what the 3000 USDT actually is A position voucher: 600 USDT of margin at 5X, BTC/USDT only, dead in 24 hours Worth taking, just know what you're taking Entry tier is 100 USDT held 3 days, the 600 one needs 2,000 held 9 days, new accounts only I'm running it with Bitunix Aug 14-27, 20 slots per round Get your slice: bitunix.com/register?vipCode…
4
11
1,831
Papailiopoulos spent 5 days telling the models their proof was unreadable The proof itself took about 30 minutes, on a MIMO detection question open since 2001, closed at the same 2 log N threshold that used to need exhaustive search His addendum matters more than the result, no new math was invented, nothing in the proof was unavailable in 2010, it just needed twenty pages of standard steps nobody would spend the time assembling Bubeck puts a clock on the same shift, mathematics went from AGI-minutes in 2024 to AGI-days by late 2025, with a COLT 2012 problem taking two days of thinking and then two more He is blunt about the ceiling too, nothing has yet met the bar for STOC or FOCS
3
14
1,969
The host opens with 193 commits in one PR he is too scared to merge It touches security, database migrations, production data and API keys at once, which is where you land when the stop condition is the agent saying it is done Alex Lavaee at Microsoft Research turned the opposite into an open source runtime called Atomic, where a claim without evidence did not happen Propose, measure with real commands like tests and type checks, then hand the result to a fresh verifier in a clean session that sees only the files and the evidence 176 minutes of him building it on camera, and the article below maps which of the three levels your failure belongs to
6
18
2,221
The compute roof is 75,000 tokens a second, and memory caps it at 187 Batch size 1 decode turns every projection into a skinny matrix-vector op, so you stream weights out of memory with almost nothing to reuse The kernel table shows it plainly, one vocab projection called a single time burns 2063 microseconds, more than 24 calls of the next kernel combined The article ends by asking what larger batches would change, and this SGLang workshop runs exactly that sweep on the same /start_profile endpoint Batch 1, 64 and 256 plotted against each other, so you can see which bottleneck survives once the GPU has real work to chew on
this is probably my favorite article ive written. in this blog, i try to profile llm inference served using sglang and reason about the patterns, kernels and bottlenecks you would usually find in production.
Article

Profiling LLM Inference with SGLang and Torch Profiler

This blog is my attempt to learn how to profile an llm served using sglang and its in built torch profiler integration that comes out of box. The idea here would be try and understand when I hit

3
2
18
2,102
Running five agents at once bills you for the same context five times Each one opens a fresh session with nothing cached, so they all rebuild the background from scratch and your plan drains 5x faster for the same work Fable 5 in Claude Code and Sol Ultra in Codex already spawn these graphs on their own, more aggressively than before, which is why plans are emptying without anyone asking for parallelism 42 minutes on when to let them do it and when to take that knob back, down to handing execution to Sonnet once the spec is already written The guide below has the 4 steps and the copy-paste graph prompt, and every diagram in it measures time, never cost
9
2
26
3,048
GitHub just showed how they actually run Copilot inference on Claude 25 minutes of production numbers most teams never get to see, starting with cache hit rate, where they hold above 94% and treat 70% as a bug, because a miss costs 10x and 1% across billions of calls is millions Three things broke it for them - a UUID sitting in the system prompt, tools loaded dynamically, and cache affinity when a session bounces from Opus to a GPT model to an OSS model and back Anthropic's Brad Abrams then demos Haiku executing with a tool that calls Opus, landing near Opus quality at Haiku prices because Opus only answers when Haiku gets stuck, and only with a hint They race both live on stage, and you watch Haiku alone spin while the advisor run finishes:
2
1
15
2,204
New campaign is live on TxFlow: Trader Royale 2.0 How it works: - Up to $3,000 per person in Fee Credits - Ranked by cumulative futures taker volume from when you join - Positions held 1 min+, manual trading only - Pool unlocks at $20M combined volume ($1,000 base), +$200 per extra $5M - Rewards paid in Fee Credits (offset trading fees, valid 14 days) Join here: app.txflow.com/campaign/trad…
9
3
17
1,422
Building the sales agent takes 5 hours, the guardrails take 8 weeks In this vid it joins a Google Meet, reads the CRM record, runs a live Tavily search, then quotes $72,000 with nobody approving the number Wayne Liang's agent took 2,741 live calls over his paternity leave, and it invented a $4,800 enterprise plan, emailed a customer its own internal urgency scores, and promised meetings off a dead Calendly link Pricing was a business rule and the agent treated it as arithmetic it could derive The article below has the three fixes and an MIT repo, though the 8 weeks went into the markdown vault the repo leaves out
5
13
2,053
you can spot an AI-built app in about two seconds the purple-to-blue gradient, the rounded cards, Inter everywhere, three buttons in the hero @boltdotnew got real design studios to build these templates instead, and every one is a working full stack app, not a theme open one, prompt it into your product, ship something with type and spacing a person actually chose all free
A truly standout website costs tens of thousands of dollars in design. So we did it for you, with a handful of the world's best designers and agencies. For free. Introducing the Bolt Template Marketplace: real, full stack apps you can open in Bolt and make yours 🧵
5
3
26
2,855
Qwen bolded itself in all 16 panels of this chart, and lost 9 of them DeepSWE never appears on the chart, and in the table underneath it reads 56.6 against Fable 5 at 70.0 and GPT-5.6 at 73.0 Next week it becomes the first Qwen Max with open weights, 2.4T total and 95B active, with a 27B alongside it 38 minutes of independent testing calls it good, but not great, after $32 of API credit One standout in the whole run, a C++ skate sim, and the tester is now waiting on the 27B instead
📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉 Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters: - Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:github.com/qwen-code-dev-bot… - Real work, real results: Production-quality deliverables across hundreds of professions. - Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy. - Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction. 💰Pricing: Input: $2.0 / M tokens Output: $6.0 / M tokens Implicit Caching: $0.25 / M tokens Start building with Qwen3.8-Max! 🚀 📖 Blog: qwen.ai/blog?id=qwen3.8 ✅ Qwen Studio: chat.qwen.ai/?models=qwen3.8… ⚡ API: qwencloud.com/models/qwen3.8…
9
1
34
2,550