I have seen many fail with crude prompts Let me teach you the hidden ways of Claude Skills, agents, and lectures

this is certified f*cking gold a 23-year-old from China runs 300 AI agents at once with Jev engineering, and none of them can lie to him he opens the dashboard live: 300 Kimi K3 agents firing in parallel Jev checking every single output against its source in about 100 milliseconds pointed at 100 companies in the EV market: > pass 1 - 12 failed: wrong revenue, dead citations, empty fields > pass 2 - 3 failed > pass 3 - zero. every row filed straight into an Obsidian vault the whole trick is the split Kimi K3 does the work, Jev decides if each claim actually matches the source Jev never writes a second opinion. it answers yes or no with a confidence score, at $0.042 per million tokens in and the loop refuses to stop while anything is still wrong an analyst team bills weeks for 100 verified company profiles. he gets them from one run one model writes, a cheaper one decides, and nothing ships until it says yes
4
3
24
3,893
Annatar.md retweeted
I STOPPED LETTING CLAUDE OPUS 5.5 MAKE DECISIONS THE DAY I BUILT THIS JEV AGENT FOLDER I used to let Claude decide, write and act on every single step -> now Claude only writes. Jev makes the calls, code does the acting, and every step leaves a receipt 10,000 decisions cost me $0.42 everything inside the folder: • the input > AGENTS.md - when to call Jev and when to skip it > state/build_state - goal, workers, done, missing, constraint. evidence, never a summary • Jev decides (questions/) > route - which model tier gets the task > next_worker - which worker moves next > relevance - keep or drop every tool output > done - is the goal really met > risky - will this send, pay or delete something? • code acts (rules/) > hard_rules - stop after ten actions, never publish anything unapproved > thresholds.yaml - act only on confident answers. fraud needs 0.95 • Claude writes (workers/) > research - sources and notes > writer - drafts and briefings > the one place in the folder where text gets generated • the proof > receipts/decisions.jsonl - options offered, chosen id, re-check, fallback > evals/ - dozens of my own labelled traces, thresholds tuned on them • the guards (hooks/) > pre_tool_use - every command checked before it runs > stop - confirms "all done" before the agent is allowed to quit median 300 ms per decision. the expensive model never waits on a yes or no again the engineer who can show this bill to their team stops being the person who uses AI and becomes the one who decides how it runs an LLM writes, Jev decides, code acts
9
24
182
18,143
Annatar.md retweeted
Jev + Claude Opus 5.5... i can't understand why everyone isn't wiring this up yet the forks alone drop from about $110 a day to $4 once jev takes them over. right now it's the best way i know to run opus what jev handles: -> scores which project notes matter before the context even loads -> hands the easy steps to a faster worker -> picks a recovery path when a tool fails, so nothing retries blind -> runs the focused checks before the full test suite opus does the hard reasoning. jev chooses from options the harness already prepared and validated that's jev engineering in one line: the expensive model thinks, the cheap one decides $106 a day back on the same work is a number your manager remembers read the paper, then the full jev guide below
59
140
1,050
125,097
Annatar.md retweeted
this is pure f*cking gold for anyone running coding agents Jev founder Diogo Amogo wrote a PDF on building a Jev harness what it claims: > 200x faster > 400x cheaper the model stopped being the bottleneck a while ago the speed and the bill both live in the harness around it • how to use it > hand this PDF and the article below to Claude Code or Codex > tell it to rebuild its own setup one evening of Jev engineering and next week your agent runs on a harness most teams haven't built yet 👇
11
41
367
111,259
Annatar.md retweeted
this is straight f*cking gold for anyone on Claude Code you already pay for Fable 5.1, and while Opus 5.5 does all the work it just sits there one command puts it on your session as a senior reviewer: /advisor fable Opus 5.5 keeps writing the code Fable 5.1 reads the whole session, every tool call included, and speaks up at three moments only: > before a plan: is this the right approach? > when the same error comes back: am I digging in the wrong place? > before "done": what did I miss? Fable 5.1 reviews. Opus 5.5 ships Jev engineering does the same thing one layer down: forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big model only sees the ones that actually split • the full tree > Opus 5.5 on high runs the main session > explorer reads the code > worker edits and runs tests > researcher pulls the docs > all three on medium > Fable 5.1 on call as the advisor paste the tree and this prompt into Claude Code ↓ "Rebuild my Claude Code setup around this tree: 1. Check ~/.claude/agents and .claude/agents for subagents that already fit explorer, worker and researcher. Draft new ones only for missing roles. Give each model: opus, effort: medium. Skip any that pin a different model and list them. 2. Set the main session to high via effortLevel in ~/.claude/settings.json, and set advisorModel to fable. 3. Find anything that keeps the advisor off (CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY, any variable that stops feature-flag fetching) plus CLAUDE_CODE_EFFORT_LEVEL, which overrides subagent effort. Report them, change nothing. 4. Add one rule to ~/.claude/CLAUDE.md: consult the advisor before a large plan, when an error repeats, and before calling a long task done. Show me every change as a diff first. No edits until I say go" ↳ code.claude.com/docs/en/advi…
19
46
398
72,739
Annatar.md retweeted
the winner of an Anthropic hackathon open sourced his full Claude Code setup - and this is FREE f*cking gold 68 subagents, 286 skills, 94 commands, MIT license ECC takes a single Claude Code assistant and staffs it like a full engineering team nothing gets built without a plan, nothing gets fixed without a failing test first, and every change is reviewed by a context that never saw it being written • the team > planning - one sentence in, a plan you sign off before any code exists > review - a clean-context read of your diff, one reviewer per language > build repair - a separate fixer for each toolchain, PyTorch and CUDA too > security - an OWASP sweep and a scanner that looks for injection holes in your agent config > architecture - flags design mistakes while they still cost minutes to undo > domain work - database queries, ML pipelines, e2e tests, docs the security pair is what almost everyone skips an outside OWASP audit costs four figures and takes a week this one is done on your branch by lunch • the skills > testing - tdd-workflow gets you from red to green, eval-harness runs over it > language packs - Python, Go, Rust, C++, Django, Laravel, Spring Boot, Next.js > context - search-first checks the docs before it writes, iterative-retrieval stops the repo from drowning the window > shipping - Docker, CI/CD, health checks, rollbacks, migrations > beyond code - your writing voice, market research, pitch decks fork it, strip it down, and run your own version tomorrow begin with one plan and one rules pack turning on all 286 skills on day one is how you make it worse with 68 subagents handing work back and forth, this is where loop engineering and Jev engineering pay off: the loop keeps the team running on its own, Jev makes the routing calls, Claude does the actual work spend one weekend wiring this in and the next quarter you're reviewing work while everyone else is still typing it
12
24
321
50,017
Annatar.md retweeted
this is certified f*cking gold for anyone paying for Claude Sonnet 5.5 costs half of what Opus 5.5 does on real office work across 44 jobs it scores 1844. Opus scores 1846 on coding it went from 10.3% to 70.6% in one version and passed Opus on the way most of what you pay Opus for, Sonnet 5.5 now does at half the price here's how to split the work: • give it to Sonnet 5.5 > bug fixes and well-scoped coding tasks > decks built from your own slide template > spreadsheets and polished documents > UI polish, it has a real eye for design > fast back-and-forth, it runs 30%+ faster • keep it for Opus 5.5 > long, open-ended work with judgment calls for hours > anything where one wrong assumption costs you a day • the setting that saves the most > Medium is the default in the Claude apps and Claude Code > on Low or Medium it beats Sonnet 5's best score for a tenth of the cost > raise effort only when the first answer falls short • the price > $2 in and $10 out, against $4 and $20 for Opus > cache reads $0.20 on both • check before you switch > if you ran Sonnet with thinking off, move to the new between_tools setting first whoever routes the weekly grunt work to Sonnet 5.5 and saves Opus for the hard calls ships more on half the budget, and that's the person who gets noticed
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
7
6
93
19,404
this is rare f*cking treasure 9 use cases where Jev can replace an LLM (TypeSafe's full list) Jev is TypeSafe AI's first System One model, 100x faster and cheaper than frontier LLMs in the right hands it opens up the work people skip because the big models are too slow or too expensive for it:
6
34
170
16,827
this is pure f*cking gold TypeSafe released the skill that teaches Claude Code to build on Jev left alone, coding agents send Jev one question per call and guess the field names this skill fixes the order before a single line gets written: docs -> behaviour -> judgments -> one request -> code decides every rule in it, broken down on two pages: > the failure behind each of its 11 instructions > Choice, Noul or Score, picked in one glance > a prompt skeleton you paste straight into your agent 12.2x that's how far the bill drops when 13 questions share one call, measured in TypeSafe's cookbook one short file, and your agent stops paying for decisions it should batch
15
14
116
10,067
this is rare f*cking gold an internal AI engineering document is making the rounds, and it's saving solo devs $300,000 a year that number is three hires you never make: > the one who writes the spec > the one who reviews the output > the one who runs the queue prompt-driven is over. loop-driven is where it's going Generate → Evaluate → Remember → Schedule → Optimize → Recurse six layers, one loop, and it improves itself with no human in the middle • the six layers > generation - writes its own brief, then builds against it > evaluation - a second layer grades the work and sends it back > memory - keeps what worked, so each mistake costs you once > scheduling - picks the next job, the queue runs itself > optimization - rewrites its own instructions from what shipped > recursion - pull out one layer and the whole thing degrades that last line is the tell. six layers is the minimum • the career part nobody says out loud > every team already has someone typing prompts all day > almost no team has the person who builds the loop that replaces that job > bring this into your company and you stop doing the work > you start designing how the work runs that's the jump from operator to architect, and it's the promotion that's actually open right now the people who wire this up over a weekend spend next year reviewing output while everyone else is still typing prompts
7
15
96
10,564
KARPATHY WROTE THIS DOCUMENT TO TURN OBSIDIAN INTO A SECOND BRAIN THAT RUNS ITSELF ON CLAUDE I was ready to abandon mine cross-referencing everything by hand was eating my week then I found this document and the whole approach flipped Karpathy's method turns the model into a full-time maintainer for your vault: • what the agent actually does > reads every new source and files it into a structured wiki > writes the links for you, so the vault compounds while you sleep > runs checks that surface contradictions between your own notes > keeps Obsidian as the visual layer while Claude works the backend • what changed for me > I feed it raw documents and stop thinking about structure > nothing gets manually filed, nothing gets lost > the friction is gone two years of notes stop being a graveyard and start being an asset you can query that's the difference between reading about your field and being the person in the room who already knows the answer here is the document from Karpathy explaining the architecture 👇
7
32
140
15,031
CLAUDE + OBSIDIAN + LOOP ENGINEERING = AN AGENT THAT LIVES INSIDE YOUR NOTES Karpathy's second brain runs on something close to this two years of growth almost none of it typed by him • the loop, four steps on repeat > read - Claude Opus 5 opens the vault, not a chat window > write - notes and links get edited inside a branch > check - a critic reads the diff and walks every link > keep - the good change lands, the rest is left where it was • the math > costs 2-4x a single prompt > pays for itself past a 5% gain > zero notes overwritten, the vault only grows • the whole system, six plain files > CLAUDE.md, skills, subagents, hooks, MCP, plugins > you can read every one of them in an afternoon • three ways in > desktop connector - three clicks > Claude Code - full control > Obsidian plugin - you never leave the app the rules that keep it safe: append, never overwrite measure the gain before you add another step don't let it rewrite the whole vault in one turn set this up on a saturday and by december your vault answers questions your team is still googling a loop that never deletes a note is compound interest on your own thinking the stack: 📁 Claude ↳ claude.ai 📁 Obsidian ↳ obsidian.md
7
18
102
11,635
Annatar.md retweeted
this is free f*cking gold a second brain article hit 8 million views, so the guy behind it put the entire setup in one place the repo, the guide, the tools, the learning path. all of it, free • the guide > 10 sections, 65 pages, concept through troubleshooting > 5 tracks on top, 44 pages, 15 of them build guides with code that runs • the machine (.claude/) > 18 agent skills, one per workflow > 72 slash commands - /ingest-pdf, /ingest-youtube, /ingest-voice, /backfill > 6 subagents - curator, linker, researcher, reviewer, ingestor, graph-analyst > 4 of them read only, so nothing rewrites your vault behind your back • the scripts (plain Python, zero dependencies) > graph export, link checker, vault stats, chat converter, site builder • the starter vault > its own CLAUDE.md with page contracts and linking rules > raw/ never edited after it lands, wiki/ is what the agent maintains > log.md - one line per run, so the whole thing stays auditable • 87 vetted resources > 28 tools, 26 Obsidian plugins, 15 repos, 12 skills, papers and articles five tracks to pick from: > knowledge graphs > Jev engineering > agent harnesses > loop engineering > eval engineering start with the second brain guide if you're new. go straight to the tracks if you already live in this stuff a consultant charges four figures to build you a research system. this one sits in a public repo under MIT ↳ github.com/undefined-ui/seco…
32
214
1,483
201,260
paste this Jev prompt into Claude and let it audit you it sets up Jev, looks at how you actually work, finds where your hours and your money leak out then rebuilds your setup around what it found same $20 subscription, and you come out ahead of 95% of the people running the same model the full prompt 👇
12
10
80
11,970