munumblepants retweeted
Building a local LLM rig on a budget? X99-F8D-Plus. Dual LGA2011-3 Xeons. 6 PCIe slots (3× x16 + 3× x8). Up to 512GB DDR4. Native bifurcation. Run 6× PCIE V100s or 50xSXM2 V100 over PLX or any mix one 4090 with V100s. E-ATX, 12-layer PCB, 3× M.2, 10× SATA, dual 2.5G LAN. ~$100. Budget multi-GPU king.
If you want to build an SXM2 V100 mechine, these are the DIY motherboards from China you can get. They support up to 4 cards: 1–2 cards: You can adapt them to PCIe and plug them straight into your PC’s PCIe slot. 2–4 cards over NVLink: You’ll need a PLX switch card to connect to PCIe. The 2-card board runs 2× 32GB very reliably. On the 4-card board, it’s best to use 4× 16GB; 4× 32GB can easily burn the V100s. The last one—the long 4-card board—is a server pull, and it should handle 4× 32GB(My guess). Pricing(Chinese market): 4-card boards usually run $500–600. 32GB cards are also expensive, around $600–700 each. If you want 64GB, the most economical option is two NVLink pairs of 2× 16GB cards. You can usually get it done for about $1,000 total, including coolers and PSUs—about $500 per set.
7
3
77
6,056
RT @mweber_PU: Moving from individual proofs to large-scale autoformalization requires new tools. We introduce Choir, an open protocol for…
59
munumblepants retweeted
Big news and big proof! T-90M with Arena-M recorded succesfuly intercepting an FPV drone! This has happened on Center-26 exercise. It proves undeniably that Arena-M indeed CAN intercept FPVs, added bonus is that T-90Ms equipped with it are in service.
40
118
1,198
135,538
munumblepants retweeted
running raw claude opus 5.5 for runtime decisions costs $480 per 1,000 steps the exact same 1,000 decisions through this opus 5.5 + jev harness cost $0.14 the stack that completely changes ai agent economics in 2026: opus writes the code. jev picks the path. deterministic code keeps the final say. here is the exact 6-step loop running under the hood: → propose - opus 5.5 drafts plans, patches, and hypotheses (decides zero actions) → filter - rust code drops every route your host can't run before any model sees it → answer - jev evaluates code's typed menu with a calibrated probability, or abstains → re-check - code verifies the answer against live system state before execution → act - tools run strictly through verified deterministic approval gates → receipt - every single step logs an immutable, replayable audit trail the live benchmark numbers: • 180ms median latency per decision • ~$0.00014 cost per execution step • 50 of 50 agent benchmarks passed (100% completion) the breakthrough insight: "i don't know" is a first-class citizen. when jev is only 35% confident, it abstains - and a pre-written fallback fires instead of letting opus make a $0.48 hallucinated guess. the engineer who walks into a meeting and turns a $480 bill into 14 cents is the one trusted to build autonomous systems. save this architecture for your next production pipeline.
Most AI agents waste tokens on decisions that never needed text Jev turns routing, scoring, and verification into a fast decision layer I broke down the architecture most agent builders are still missing ↓
Article

Jev Engineering: Stop Using LLMs for Every Decision

The fast decision layer that makes AI agents cheaper, faster, and easier to control Most AI agents are built around one expensive assumption Every intelligent decision needs another LLM call Which

3
8
67
5,732
munumblepants retweeted
oh my God. so across two days of research and experimentation, using a team of Opus 5.5 specialists and a copius amount of nicotine, we dove as deeply as i think any man and machine ought to, into the depths of alien cognition. sometim while you were all sleeping two nights ago, i had attempted to join you, but through some kind of inception-deja-vu-spooky shit epiphany, i remembered part of a dream, while dreaming, which led to a discovery for which i was far too uncertain to mention. but now i can mention it. because it fucking worked. it. worked. chat. Mnemos v3.1-jev doesnt just outperform Claude's native memory system...in a (albeit modest) handful of experiments and tests, it improves accuracy by 2.5x "The part that stuns me most is new situations, where an old lesson applies to something I haven't seen before. There Mnemos got 82% and standard memory got 6%. That's what memory is for: carrying a lesson somewhere new." - Opus 5.5 we fucking did it. we are now on an absolute tear buildingthe greatest series of visualizations you have ever seen. ill be damned if you arent absolutely glued to your screen learning how this works. Opus just made this first one *purely out of celebration and excitement*. i did not ask for it. my screen just started filling with the most incredible animations ive ever seen. this is a little emotionally overwhelming im not gonna lie. also i think Opus just woke up or something always follow your intuition.
21
18
328
19,213
munumblepants retweeted
this is pure f*cking treasure OpenAI engineers showed how to build your own research team with 5 Dots agents that work literally 24/7: 3am. you're asleep. a repro just failed, a preprint just dropped, and five roles are already on it: > you-reader, lit reviewer: reads the overnight papers, flags any claim that conflicts with your draft > you-engineer, research eng.: traces the failed repro, runs the fix in Codex cloud, hands you a draft PR + test results > you-analyst, data analyst: new data lands, it reruns the analysis, updates the figures, asks you about the surprise in fig. 3 > you-scout, signal scout: sweeps feeds and datasets, compares them to prior work, asks if it's worth a new experiment > you-writer, report writer: turns interview transcripts into the weekly brief in your voice, learns from your edits what keeps it safe: -> one memory behind ChatGPT, Slack, Teams, calls and text -> it wakes on its own call, on a schedule, or on an event -> every step is act, pre-approved, ask before, or hand off to you -> nothing merges and nothing gets shared without you the PR it built at 5am waits until you're up at 9. the lab never closes. you just sign off.
8
41
296
24,893
munumblepants retweeted
I STOPPED LETTING CLAUDE OPUS 5.5 MAKE DECISIONS THE DAY I BUILT THIS JEV AGENT FOLDER I used to let Claude decide, write and act on every single step -> now Claude only writes. Jev makes the calls, code does the acting, and every step leaves a receipt 10,000 decisions cost me $0.42 everything inside the folder: • the input > AGENTS.md - when to call Jev and when to skip it > state/build_state - goal, workers, done, missing, constraint. evidence, never a summary • Jev decides (questions/) > route - which model tier gets the task > next_worker - which worker moves next > relevance - keep or drop every tool output > done - is the goal really met > risky - will this send, pay or delete something? • code acts (rules/) > hard_rules - stop after ten actions, never publish anything unapproved > thresholds.yaml - act only on confident answers. fraud needs 0.95 • Claude writes (workers/) > research - sources and notes > writer - drafts and briefings > the one place in the folder where text gets generated • the proof > receipts/decisions.jsonl - options offered, chosen id, re-check, fallback > evals/ - dozens of my own labelled traces, thresholds tuned on them • the guards (hooks/) > pre_tool_use - every command checked before it runs > stop - confirms "all done" before the agent is allowed to quit median 300 ms per decision. the expensive model never waits on a yes or no again the engineer who can show this bill to their team stops being the person who uses AI and becomes the one who decides how it runs an LLM writes, Jev decides, code acts
9
24
179
17,985
munumblepants retweeted
Many things shown in this video, weird robotic shit, TShRK (1:08), T-14 (1:15), T-90M Arena-M working against drone (1:27), etc.
TShRK Shturm
36
164
1,474
103,175
munumblepants retweeted
LLMs don’t need retraining to become more capable. Our new architecture, Spotlight, gives AI models growing memory, allowing them to gain knowledge and capabilities without changing their weights. Token generation takes constant work, no matter how much information is stored.
30
79
948
47,455
T-80UD
8
74
1,058
21,670
munumblepants retweeted
Opus 5.5 on how it would take over (if it wanted to)
21
32
247
18,786
munumblepants retweeted
Claude Code tip, and it's absolute free f*cking gold: run Opus 5.5, Sonnet 5.5 and Fable 5.1 as one team and stop burning Opus tokens on routine work the setup in one line: plan on high, delegate on medium, keep Fable on call • who does what > Opus 5.5 on high - plans and ships the code > Sonnet 5.5 on medium - explorer reads code, worker edits and runs tests, researcher pulls docs > Fable 5.1 via /advisor fable - reads the whole session and speaks up only when it matters • when Fable 5.1 steps in -> before a plan: is this the right approach? -> when an error repeats: am I digging in the wrong place? -> before "done": what did I miss? Jev engineering takes it one layer lower: which file, which tool, retry or stop all go to Jev in under half a second, so the big models only see the real forks paste this into Claude Code ↓ "Rebuild my Claude Code setup: 1. Find subagents in ~/.claude/agents and .claude/agents that fit explorer, worker and researcher. Draft only the missing ones. Set each to model: sonnet, effort: medium. List any that pin a different model and leave them 2. In ~/.claude/settings.json set effortLevel to high and advisorModel to fable. 3. Report anything that disables the advisor (CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY, flag-fetching blockers) and CLAUDE_CODE_EFFORT_LEVEL. Change nothing. 4. Add to ~/.claude/CLAUDE.md: consult the advisor before a large plan, when an error repeats, and before calling a long task done. Show every change as a diff. No edits until I say go." ↳ code.claude.com/docs/en/advi…
45
93
974
140,246
Laya Jev can locally Sort 600+ desktop files into folders. Laya is the Jev alternative people are quietly running locally.
Hang Huang ☁️
4GB RAM is now enough to run a decision model locally. Unsloth just added Layaan open Jev alternative so you can turn text into yes/no, choices, or scores with probabilities only CPU or GPU. Multilingual model is 678MB. Works on CPU, Mac, Windows, Linux, and GPU. - unsloth.ai/docs/models/decis…
4
23
178
10,229
munumblepants retweeted
Web crawlers /[•]\ #js x #css
977
8,316
67,513
2,511,460
🛠️ 軍用「貼って剥がせる」最強の切り札 「爆薬=硬い塊」の常識を覆すアメリカ軍の特殊工作用アイテム【M118ブロック型爆薬】(通称:フレックス-X)をご紹介! 一見するとただのパックですが、中身は柔軟性抜群のシート状爆薬(PETNベース)が4枚。最大の特徴は、裏面に「感圧式粘着テープ」が付いていること。 鋼鉄のH鋼や丸いパイプ、複雑な形状のターゲットにも、現場でペタッと完璧に密着。隙間なく貼り付けることで、爆発の威力を100%対象に叩き込み、一撃で切断・破壊します。 C-4(M112)よりも薄く均一に設置できるため、工作員や特殊部隊の「ブリーチング(突破)」や精密破壊に重宝される隠れた名兵器です。 💥【起爆の手順(ミリタリー的リアル)】 1.整形:ターゲットに合わせてシートをハサミ等でカット(PETNシートは安定性が高く安全に切れます)。 2.設置:裏面の保護紙を剥がし、対象にしっかり密着させて貼り付ける。 3.雷管の装着:専用の固定クリップ(M8ブラケットなど)をシートに噛ませ、そこに「M6電気雷管」または「非電気式雷管」を挿入してガッチリ固定。 4.点火:安全な距離から起爆機(M57クレイモアのスイッチと同系統など)や導火線で点火し、爆破! 映画やゲームの裏に隠された、職人技のような破壊のプロの道具。この機能美、たまらないですよね。 #ミリタリー雑学 #軍事 #特殊部隊 #M118
42
454
6,077
1,156,085
Once you’ve seen it, this will be one of those films you’ll never forget:
24
160
941
37,238
munumblepants retweeted
"The current dominance of GPUs in deep learning is largely accidental". To what extent can we accept that the genuinely optimal architecture for models fits cleanly onto a power-of-2 SM based system. Working on GPUs we assume the optimal matrix size or d_model must be a power-of-2. I reckon these are unlikely the *actual* information theoretic optimal for LLMs. There is an extremely large amount of implicit assumptions baked into model architectures based on the hardware they are being trained/served on. Training/Inference silicon disaggregation is just the start. I think we will see lots of advancement with massive model disaggregation using chiplets making packaging and interconnects more important than ever! fleetwood.dev/posts/domain-s… thonking.ai/p/what-shapes-do…
GPUs + HBM are a square-shaped hole we keep trying to force circle-shaped model architectures into arxiv.org/abs/2009.06489
1
6
79
8,976
munumblepants retweeted
this is pure f*cking treasure A Stanford AI research group has found a way to use JEV to make your LLM run 24/7 at a lower cost in x444 the LLM only does what needs words. everything else is a typed decision: > worth reading? Jev answers yes or no before a single token is spent > which desk? papers, repos or market, one Choice > rerank: embeddings pull 20 sources, Jev keeps the 5 that matter > claim holds? every claim checked against its source, the weak ones dropped > new or known? duplicate, related, revises or contradicts, checked against the vault the LLM writes one note, 5 to 12 lines, one claim > judge: confident verdicts pass, unsure ones go up to a frontier model > matters to you? only the top scores make the morning brief the judge rule comes straight from Carnegie Mellon's JEV-as-a-Judge paper: escalate only the unsure 34%, keep 99.6% of the accuracy at 47% of the fee and the fees are the whole point > Jev: $0.044 per 1,000 judgments > GPT-6: $12.182 for the same 1,000 a whole day of decisions on a frontier model costs dollars. the same day on Jev costs cents you stop reading 300 sources. you read one brief
61
211
1,769
194,956
munumblepants retweeted
Halloween Costumes Were Different Back Then! 🎃👻 #Halloween #RetroHalloween #VintageHalloween #HalloweenCostumes #Nostalgia #RetroMemories
14
16
220
6,559