🔬 MD fueled by a deep passion for Medicine & Tech. 🌐 Exploring the frontiers of VR/AR/MR & BCI. 🤖 AI enthusiast.

Pinned Tweet
The future may belong not to those who know the most, but to those whose identities can survive being transformed by what they learn.
6
2
35
5,772
Agreed
Buying 4 DGX Sparks was a INVESTMENT... A pretty HUGE one. But the way these prices are moving I feel like I dodged a bullet. GLM 5.3 Flash on 4 x DGX Sparks is very fast, I think I will thank my past self alot in the future.
1
50
TechMD retweeted
나도 Spark 2대를 사면서 이 구매는 지금 시점에서는 정당화 할 수 없다. 클라우드 비용이 더 싸서 5년 써도 회수 못한다. 그러나 지금 ChatGPT 플랜은 지나치게 싼 미끼상품이고 가격이 오를게 확실하다. 이런 생각으로 리스크 헷지를 위한 투자상품이라 생각하고 구입했습니다. 이후 어떻게 되었냐면 1) GX10괴 DGX Spark 가격이 50% 인상 됐습니다. 2) 램이 절반으로 줄어든 Spark가 128GB 구입 가격보다 비싸졌습니다. 3) 텐서폴드의 등장으로 GLM 5.3 Flash가 60tok/s로 일합니다. 4) OpenAI는 Pro 20x 플랜의 제공량을 절반으로 줄였습니다. 투자로서는 아주 성공적임이 틀림 없습니다.
Buying 4 DGX Sparks was a INVESTMENT... A pretty HUGE one. But the way these prices are moving I feel like I dodged a bullet. GLM 5.3 Flash on 4 x DGX Sparks is very fast, I think I will thank my past self alot in the future.
1
2
26
1,745
Amazing set up
2
155
TechMD retweeted
Ring of four Sparks 1 -> 2 -> 3 -> 4 -> 1 Any recommendations on what to try first?
1
1
6
335
TechMD retweeted
🚀🚀🚀 GLM-5.3 Full CAN SEE NOW 👁️👁️👁️ The full model. On 4× DGX Spark. In your house. It was text only this morning. Inspired by @MiaAI_lab and the great open-source community github.com/Leoid/GLM-5.3-Ful…
2
3
13
1,119
TechMD retweeted
oh wow GLM-5.3 Full on TP=4 DGX sparks .. that looks so good....@WescheNex1q thank you for testing this.. want to let you know the recipe speed is now improved even more.. wait for the next update soon.. thanks bro!
GLM 5.3-753B vs Qwen3.8-27B Same Eiffel Tower prompt, both local -GLM-5.3 full (4× DGX Spark): 120.9K tokens -Qwen3.8-27B BF16 (Mac Studio): 78.5K tokens The 28× bigger model thought 54% longer and still finished in half the time. Look at the towers. Setup: GLM-5.3 = @u1tra_instinct EXL3 2.75bpw + TensorFold TP4 recipe on 4× DGX Spark (MTP drafts, 262K context). Qwen3.8-27B = official BF16 via MLX. Temp 0.7, thinking at max, no output cap. Both MP4s checked: 1200 frames, 20.000 s, 1080p60. Fixes logged in REPAIRS.diff; I can share the method and scripts.
3
1
13
542
Dreams money can buy. Wow 🤯 5 Pro 6000 Max Q set up.
できた。GPU6枚認識しました。 GLM5.3 FlashのAblitratedに、GPUを使ってAblitrated作業を自律して行いように指示しておいた。 つまり、ローカルLLMのチューニングをローカルLLMにやらせてます。計算資源が行き渡る未来では当たり前の光景かもしれないね。
2
17
845
4 in 1 remote KVM Gli Comet X is here excited
4
1
13
647
What’s the best hardware deals for Prime day? @amazon Comment below 👇
2
2
733
Next level 540 tks on Deepseek
V41 Flash 6000 TP4 539 tok/s code recorded, 518 tok/s median +350core/+6000mem stable OC Now that I got your attention, we as community really need to come up with a standard measurement for decode/prefill, and start questioning any numbers people put out including mine. Too often I see a headline, actually check their work to find it’s inflated 20-30% on top of using extremely predictable prompt can add +100 tok/s with drafters. Most people do not have 4 6000s to confirm these numbers, so it simply goes unnoticed and it bothers me greatly that they are preying on communities trust to simple believe what is being shown I’d like to think local AI is all about trust and we should always do our due diligence when we share/receive data in this community. I personally will always spend the extra effort to measure every single data I put on here, be as transparent as humanly possible, and freely admit if I’m ever wrong about anything without hesitation to gain your trust. Even then you should always question my work and numbers, and demand nothing but absolute honesty out of everything you see in this community so people can’t get away with shady business. Anyways, good morning and happy Sunday
1
6
717
TechMD retweeted
Spark Doctor v0.3.1 is out for DGX Spark 🩺 A few fixes and new checks in this one: • Fixed the v0.3.0 bug that kept reporting scans as “Incomplete” on real Sparks • Added checks for kernel OOM kills and NVIDIA Xid errors • Fixed false low power alerts during memory bound decode, thanks @scottleimroth! github.com/joeynyc/spark-doc…
6
8
76
2,705
TechMD retweeted
Anytime now 👀
Qwen 4 27B.
6
2
77
4,990
Sonys (dumped) PSVR mascot is now on PCVR and Quest! and the PC version is absolutely next level. Super high fidelity and super stunning. This is how it was meant to played👉piped.video/SY-KoLUwCBc?is=yXPI…
14
6
94
3,386
TechMD retweeted
Replying to @argofowl
Heard
52
5
533
49,703
TechMD retweeted
🚨🚨Fastest Mac/Apple Silicone MiniMax H3 via TensorFold just dropped🚨🚨 the repo is so good.. i said fuck it just drop it let other people play and test......(can be more refined and optimized) + big ass credit section for all the mac pioneers from MLX development, antirez, Ash hart.. too many to list.. please follow the original creators Apple Silicone/Mac users gonna eat good on MiniMax H3 video generation using TensorFold , PR submitted to Tensorfold Github Github Repo/Pre-built image here: github.com/drowzeys/keys-Mac…
9
10
93
4,972
TechMD retweeted
From 1x3090 to 20 DGX Sparks: my house fuses were the first bottleneck.. From the first LLaMA 33B I knew I wanted that magic-like intelligence locally, mine, so nobody could take it away when I needed it. I bought a 3090 for my home PC. Then LLaMA 65B appeared and I was dazzled, it looked like it had all the knowledge in the world. I made two copies, one local and one on my Synology NAS RAID, so I'd never lose it, and bought a second 3090 to run it. I was happy for a year with small coding tasks on LLaMA and Qwen models. Then DeepSeek 671B MoE appeared. Wow, frontier level at home. I upgraded to a Threadripper with 512GB DDR4 and ran it at 8 t/s with experts offloaded to RAM, or Qwen 235B at 10-12 t/s when I wanted speed. I used these for real coding at my job, in OpenWebUI. Then agentic coding took off and this was too slow. At 100k context generation speed halved and prefill made it a beautiful yet agonising experience. So: 16x3090 across P620-based nodes on a 100Gbit network. It ran MiniMax M2, Qwen 235B and even Qwen 397B, as good as anyone could desire. I built an entire paid project with 397B in OpenCode. But bigger models were out of reach, and the house circuit said no: the fuses blew whenever the rig and the electric oven ran together. Heat and stability were issues too. Next came 4x ASUS GB10, after I read they can be linked (3 was the biggest supported config). 397B at 30 t/s on 400W, versus 50-60 t/s at 6kW, rock solid and almost silent. A dream come true. I built two more projects with it. Then MiMo 2.5 Pro and Kimi 2.6 appeared, smarter and more productive. I found no published solution for an 8-node cluster, but I still bought four more GB10s and made it work. 397B ran at FP8 instead of INT4, and 20% faster. I posted the first MiMo 2.5 Pro and Kimi 2.6 solutions on 8xSparks on the NVIDIA forum. I liked the result so much that I talked my older brother into buying his own 8x GB10, so he could run the best open models locally too, in privacy, without depending on API availability and rising costs. His house is a 5-minute walk from mine. When Kimi K3 (2.8T) appeared, biggest and smartes open weights model, we joined the clusters: two 8x clusters for daily use, or one 16x when we want the biggest model at home. After some work I published the first working solution for Kimi K3 on 16x Sparks on the NVIDIA forum. Through multiple iterations, it went from an unusable 7 t/s at 100k context to a fairly usable 20 t/s at 300k. Now we're adding 4 more Sparks, so a smaller, faster model (GLM 5.3 Flash) runs 24/7 while the big cluster runs either GLM 5.3 on 8x plus MiMo 2.6 Pro on the other 8x, or 16x Kimi K3, or Qwen 3.8 2.4T. I'm always tuning speed on the big models and rebuilding vLLM/SGLang images, so always-on smaller cluster made sense, why? Because for all my work projects and my vllm/sglang personal projects, I chose to use only local hosted models, I never paid a comercial model subscription, not because of the cost, but, because of my strong confidence in local models future. They arrived October 2, along with 4 more Sparks for my younger brother, who got caught by the same local AI microbe :)
13
11
81
7,826
TechMD retweeted
🚨 | NEW: Malaysia's Prime Minister has confirmed the government are currently in talks with Formula 1 about Sepang returning to the calendar. [@straits_times]
131
1,621
32,042
407,364
TechMD retweeted
GLM 5.3-flash vs Mimo v2.6 Round 2 of the local Eiffel Tower duel. GLM 5.3 Flash: 80K tokens in 61 min. It needed 3 one-line fixes to run, then delivered the best tower of the four: real cross-braced lattice and base arches. MiMo V2.6 Flash: fastest (62.7K tokens in 23 min) and clean code first try, but the low shot falls through the ground into a black silhouette.
Deepseek v4.1 vs Qwen-next Same Eiffel Tower prompt, two local models on 2 DGX Sparks, thinking on. DeepSeek V4.1 Flash wrote 114K tokens in 49 min and the code ran first try: real cross-braced lattice, the Seine bend, La Défense in the haze. Qwen3.8 Flash Next was faster (77K tokens in 21 min) but needed 3 code fixes to run, and the river never shows.
3
1
28
2,469
TechMD retweeted
Some people have asked about my local hardware setup My local engineering team runs Qwen3.8 Flash Next on a 256 GB M3 Ultra Mac Studio and GLM 5.3 Flash across two HP ZGX Nano G1n machines, linked through MCDMA. They work in OMP with Duo and my experimental Drift link, which exchanges translated KV context between the local models. For frontier work, I use Opus 5.5 as my CoT and guide GPT-6.1 Sol through tmux along with my two local models to review the code and tests, keeping most of the engineering local and cutting my cloud token usage. github.com/ashhart/MCDMA github.com/ashhart/Drift github.com/ashhart/Duo github.com/ashhart/TensorFol…
12
10
157
6,533