| Just another Army Grunt | Delivering decentralization @hlabs_tech | Spectrum: Decentralization.

Decentralized Foxhole
Pinned Tweet
Local inference, be proficient:).
Made with AI
1
6
157
I almost forgot this huge recent update... The CIP-113 tech for programmable tokens is one of the ultimate milestones to get full Layer Zero functionality with Cardano. So much cool stuff going on 😅😃
Congrats to the many authors, contributors, and teams that have spent literal years making CIP-113 a reality! It's finally merged as its next step towards a path to "Active" status! 🎉🎉🎉 @MicheleHarmonic @phil_uplc @0xMetamatt @CryptoJoe101 github.com/cardano-foundatio…
2
6
74
1,362
cap.intersectmbo.org/#/detai… The Dijkstra hard fork introduces new updatable protocol parameters for reference-script protection, Ouroboros Leios, Ouroboros Peras and pool economics (CIP-0023 minimum pool margin, CIP-0050 maximum pledge leverage)…
1
1
7
215
Mike retweeted
Use this rotation movement
Fight Club
1
84
557
13,640
AMD basically caught up. Nvidia freaked out and released a 64Gb Spark 😂😂.
I tested the best 3.8 flash models against each other on the Strix Halo. Then I decided to test it against Qwens token plan API. AMD is faster at real work. llm.ciru.ai/research/flash-r… For more comprehensive benchmark data see quoted tweet.
1
2
189
Some people buy a 64Gb Spark and some 4xRTX 6000 Blackwells 😅.
主役が届きました。RTX Pro 6000 Blackwell Max-Q x4 これで今年の冬は暖房が不要になりました。
87
Mike retweeted
I am done being polite about this industry. Not disappointed. Done. I spent the last months pulling wallets apart instead of posting charts, and what sits in my database now is not a pile of unlucky trades. It is hundreds of thousands of connections between wallets, coins, launch tools and payment flows, and they draw the same picture every time: industrial draining. On Solana first, because that is where the machine runs fastest, but the same hands show up on other chains under other tickers. This is not degens getting rekt. This is organised crime with a Pumpfun profile and a leaderboard rank. You want to know what "KOL" means in 2026? Last week I followed 3 wallets off a copy trading leaderboard. Their money does not come from calls. It comes from wallets that create coins by the hundred and by the thousand. One of those feeder wallets has 4,578 coins to its name. In September alone 852 SOL net moved from the factories to the 3 "influencers". One coin was created, dumped 33 seconds later, and the money sat in the KOL wallet 2 minutes after that. Then I checked the big names. Leaderboard wallets with 103,000, 134,000 and 146,000 followers on Pumpfun that have personally created 1,759, 953 and 1,217 coins. One of them is at 6,464. Create, buy in the same transaction, sell everything 3 to 5 blocks later into whoever got there in that second. Same second, or the next one. Over and over. And these people get called traders. They get podcasts. They get "gm". And the platforms? Fee on the way in, fee on the way out, and they call it volume. So no, I am not writing you another post-mortem every time retail gets emptied. The post-mortem is always the same and it changes nothing, because the people who benefit are the same people who get the reach. Every "isolated incident" conveniently skips the question of who keeps getting paid, who keeps promoting, and why the exact same wallets are back in the game before the thread has finished loading. I will keep saying it, because this space forgets everything the moment there is a new bag to chase. Some of my posts will get fewer views for that. Fine. But anger alone is content, and content is exactly what they want from me. What I want is tools. You should be able to look at a wallet and see its 1,217 launches before you ape. You should see who funded the dev, which "KOL" he paid last week, and what you are about to sign, before the damage is done, not in my thread 2 days later. That is where MASTR goes now. The intelligence machine gets wired into my apps and my sites, so the research lives where you trade instead of dying in a timeline. The people who support this work get that intelligence first. That is the only alpha I care about: knowing what is happening before someone else's exit becomes your loss. And I will say the uncomfortable part out loud: I cannot run this for free forever. Data, infrastructure, development, investigation, it all costs real money, even when the finished tool makes it look effortless. I refuse to pay for it by renting my name to whichever project needs a clean reputation this week. I also refuse to pretend I can subsidise everyone out of my own pocket while the drainers cash out your SOL and climb the next leaderboard with it. If I survive in this industry, it will be because I built something useful enough that people choose to back it, and stayed angry enough to keep pointing at the machine it runs inside. Receipts are coming. Every transaction linked. Check me.
93
125
633
20,331
Same Mac. Same model. Same words. My 512 GB M3 Ultra ran GLM-5.3-Flash on MLX last week: 27 tok/s, first token in 0.46 s. Today it runs TensorFold 0.6.2: 61 tok/s, first token in 0.20 s. In the clip it laps the reply twice while the MLX side is still typing. Same reply, byte for byte. Then eight requests at once, 78 tok/s combined, every answer identical to its solo run. The old engine took one at a time. Context stays at 1,048,576 on the Mac. Every number measured on my own machine with the same probe, one week apart. Nothing estimated. @ashxhart built it. tensorfold.dev Apache 2.0.
14
8
133
27,318
Really struggles
Only software engineers will understand why this is actually funny.
3
135
Nvidia Rug Pull 😂.
54
Mike retweeted
No one dreams of becoming a 64GB Sparker. Snap up the last of the 128GBs before NVIDIA turns them into a collector’s item. Where my 128GB Sparklers at? ⚡️⚡️⚡️
19
1
65
2,779
Mike retweeted
I have been working on a new inference engine that is solely focused on AMD hardware. As patching vLLM was never a long term solution. This is still a work in progress and realistically weeks away from release still. Here are the prefill results from BetterBench on my dual R9700's for Unsloth's Qwen3.8 27b nvfp4 . #AMDAI @AIatAMD #R9700 #LocalLLM #LocalAI
19
12
109
4,655
Nvidia wants you to think this 64Gb Spark for $4999 is gods gift to Ai. When the 128Gb version is already slow as shit. For that money get either the apple M5 mac mini or buy some GPUs. 64Gb is going to leave you crying yourself to sleep every night.
62
Local AI is looking like BTC mining 10 years ago 😁. I love it!
The EX8796 PCIe Gen3 Switch EVM is a Broadcom/PLX switch board that lets you hang 10 V100s off a single host. Up to 320GB VRAM (10× 32GB) or 160GB (10× 16GB). A full 10× 16GB setup runs about $1,000 total—roughly $100/card. For local LLM inference, that's unbeatable VRAM-per-dollar. No server chassis needed.
1
3
206
Mike retweeted
We are going to prove that the eUTXO model can have a flourishing DeFi ecosystem.
7
5
65
1,432
Mike retweeted
Oh boy oh boy... Now you NVIDIA FANS can have 64GB for the price of the original 128gb GB10. No spec bump. Just less memory. Cuda cores or something
24
3
57
4,801
These all in one AI boxes look really sweet and are very tempting, especially when the phrase "High bandwidth memory" is thrown into the mix and you see numbers like 128Gb to 194Gb of unified memory. And that's awesome but what they don't tell you is that there is nothing high about that memory bandwidth, except for the price tag. Yes they run some impressive open weight models, but your basically waiting for Grandma to finish crossing the road(no offense grandma). Even Apple is putting them to shame.
Do not pay $12,000 for 2 x DGX Sparks. For the same price you can buy 4 x 170HXs and a computer to run them in, for FAR better performance. Should I post complete builds at various price points?
1
141
Mike retweeted
A fine-tuned model can outperform frontier models and be cheaper and faster to run. Literally, every company I've met wants this. I want you to see these results from fine-tuning Qwen3 4B on AWS. It smokes both the out-of-the-box model and Claude Sonnet 4.6.
66
155
1,551
84,564
The AMD support is AMAZING!
- Prompt processing on multi gpu +20% - AMD performance increased by 12% - Unsloth UD-Q4_K_XL experimental fix - OrcaRouter IQ3_XXS added - RTX 20 series fixes - Many more improvements github.com/Niko1221/Strata/r…
2
220
Exactly this. I was actually mind blown when I figured out the Gotgon Halo has literally 1/5th of the memory bandwidth the M5 does. Literally a machine purposed for AI is getting it's ass handed to it by a multi purpose desktop studio device 😅. If it was $4000-$5000 for the Gorgon then I could maybe justify the hardware but for less than $3000 more @ around $10k USD you can get a beast that has almost 5x the hardware and compute power.
Hate to say it here, but the new Gorgon Halo at it's price today is NOT the value the Strix Halo offered. I configured a M5 Ultra and a Framework AI Max 400 both w/ 2TB storage. Hate to say it, but on raw $/GB the two are a wash; on bandwidth-adjusted $/GB the Mac is ~3x cheaper. System $/GB │Mac $39.06 | FW $38.07 Memory bandwidth | Mac 1.2 TB/s | FW 273 GB/s Bandwidth-adjusted $/GB ($ ÷ GB × TB/s) Mac $32.60 │ FW $140.00 the Mac has 1.33x the memory and 4.4x the bandwidth for only 1.37x the price while the framework looks $2,690 cheaper, you are not actually getting the same level of value for your investment returned to you. And since thunderbolt & usb4.0 are pretty dang fast, you can shave another $500 off the Studio by dropping down to a 1TB storage drive and using external storage if you got it. For local AI.... Apple is starting to look like they want to take NVIDIA and the DGX Spark head on. It's a HELL of a value.
1
2
312