Software Engineer shipping real software with AI. Local LLM experiments on 2x DGX Sparks. Build Things on YouTube. Mentor to 1,200+ AI builders.

Bangkok, Thailand
Pinned Tweet
7 hours ago, I wasn't confident prose would move. It moved. 1.26x → 1.36x, up 8-10%. DeepSeek 4.1 Flash on TensorFold is Live! (Still some bugs to resolve) Same day, G13 build on 2x DGX Spark: prose 40 → 44.3 tok/s (32.5) single stream 83 → 85.1 tok/s (kit 32.2) 4 streams 93.5 → 96.6 (kit 37.6) code 79-81 → 82.8 (kit 42-45) 5 rewrites, 2 were faster on the bench and slower inside the CUDA graph so they went in the bin but stayed in the docs. Fused MoE chain saved 9-46us a layer on its own and cost +0.7ms in the graph. Quality hasn't moved: 99.6% agreement, MMLU 87.5%. Still a little more on the table for this Anyone else had a kernel win in isolation and lose in the graph?
8
28
1,913
Just casually rented some B300 servers... I think we can go further, and I must have the compute! What do you guys want to see most? Different quants? Fine-tuning? Something weird?
1
7
408
Yup, I have 2 Claude Max subs now, doing around 2-3B tokens per day of real work. It's not uncommon for me to have 10-15 agents running at once, and each agent will have 2-12 subagents. I only manage 4-5 of those agents. They delegate to worker agents.
TIP: Split a task into different pieces and run them in parallel. Let each agent focus on one thing. That's the fastest way to finish a big task.
3
4
633
Yesterday I dropped DeekSeek on @ashxhart TensorFold when I crossed 400 followers. Then @MiaAI_lab gave me a shoutout in her post! Thanks, Mia! Now I am cooking something even bigger for 1000 followers! Let's GO!
When I get to 400 followers, I will share something very cool!!! ;) Hmmm, how long will it take?
1
25
1,044
DeepSeek V4.1 Flash on 2x DGX Spark: decode is 2.6x the standard kit now, with exact output. Seeing more gains across the board, but someone reported a memory leak. So the speed work is on hold. Hour-long soak tests until memory stays flat, then the recipe update. MY goal is usable recipes, then optimise. Has anyone else experienced any issues?
8
2
41
2,057
This is the strangest argument. 1 guy knows nothing about software engineering and spends more on AI subscriptions in 1 month than I do in a year. He then turns around and slanders the company the moment their latest model isn't the best offering, only to buy 5 new subs 2 weeks later. The other is a highly respected engineer who is somehow focused on the wrong thing and should be focusing on innovation with his background rather than debating nonsense. Never would I expect that this is where tech would go. Everyone is working hard to push us further. Please just offer insights when you actually have value to share and stop talking when you don't have something nice to say.
@theo, it's very clear you're jealous of BridgeMind. Some of your questions were fair. I'll be sharing more on the NerfBench methodology soon. But asking me to take NerfBench down and post an apology written by you isn't about better benchmarks. It's about your ego, and not wanting anyone else to exist in this space. More benches are good for everyone. Build yours.
2
4
700
7 hours ago, I wasn't confident prose would move. It moved. 1.26x → 1.36x, up 8-10%. DeepSeek 4.1 Flash on TensorFold is Live! (Still some bugs to resolve) Same day, G13 build on 2x DGX Spark: prose 40 → 44.3 tok/s (32.5) single stream 83 → 85.1 tok/s (kit 32.2) 4 streams 93.5 → 96.6 (kit 37.6) code 79-81 → 82.8 (kit 42-45) 5 rewrites, 2 were faster on the bench and slower inside the CUDA graph so they went in the bin but stayed in the docs. Fused MoE chain saved 9-46us a layer on its own and cost +0.7ms in the graph. Quality hasn't moved: 99.6% agreement, MMLU 87.5%. Still a little more on the table for this Anyone else had a kernel win in isolation and lose in the graph?
8
28
1,913
@ashxhart More PR's incoming. Sorry not sorry haha
1
103
Recipe + full tables: [github.com/jayleaton/deepsee…] Yes, 2-stream is still terrible. Looking into it. Also added a strict mode with every precision shortcut off. Still 2.45x single stream. The kit beats it on cold prefill, though (915-965 vs 1,031-1,075), so that's next. The 2 losing rewrites are still in the engine, set to 0, if anyone wants to dig in.
2
143
DeepSeek 4.1 Flash on TensorFold, 2x DGX Spark. Early WIP version. Across-the-board average is 2.2x faster Tonight's build vs vLLM kit from @MiaAI_lab Single stream: 83 vs 32.2 tok/s (2.6x) 4 streams: 93.5 vs 37.6 (2.5x) Code: 79-81 vs 42-45 (1.9x) Prompt Processing: 1833-2069 vs 1070 (1.8-2x) Same answers, just faster: 99.6% agreement with the kit, MMLU 87.5% (the original kit's own score), structured output 22/22, 30-minute soak with zero errors. Prose is the one that won't move yet, only 1.26x. More on why below. I am working on known bugs, but you can see the recipe in the comments. (WIP, and more updates to come.)
10
7
78
5,816
github.com/jayleaton/deepsee… Early WIP. There are known issues I am working on, but I have to prioritise compute. Hopefully @ashxhart can get my interface layer merged in which will make very easy to build out many different model layers without giving up isolated performance. :)
3
373
Replying to @MiaAI_lab
Weakness: - Prose = 1.2-1.4x - 2x Concurrent Stream - 300k context x 4 Other gains. - Less than 1 min boot time - NVMe cache offloading like GLM - structured is close to 3x
3
436
When I get to 400 followers, I will share something very cool!!! ;) Hmmm, how long will it take?
1
13
1,602
MiniMax H3 video on a 16 GB RTX 5070 Ti and 64GB memory running windows: 623 seconds per 5-second clip with stock ComfyUI. 81s with TensorFold. 1344×768 with sound. You swap one loader node and keep the rest of your workflow, LoRAs included. stock, 20 steps: 623s TensorFold, 20 steps: 273s (2.3x) + sparse attention: 192s (3.2x) + turbo LoRA, 8 steps: 133s (4.7x) + turbo + sparse: 96s (6.5x) + int8 video VAE: 81s (7.7x) For scale, people report ~133s for a similar turbo clip on a stock 4090. Free recipe, workflows included: github.com/jayleaton/minimax…
6
10
957
What changed under the hood: - the 19B core runs as 4-bit NVFP4 on TensorFold's kernels, 21 GB of weights down to 10 GB - 8-bit attention, 23.9s to 8.6s per step (attention is two-thirds of the work) - whatever doesn't fit in VRAM streams from system RAM, with no measurable cost - no more 20 GB model reload before every decode - ~600 GPU stalls removed per step, fused kernels for the small ops Will also add a DGX Spark version when my Sparks have some downtime.
2
196
Qwen-Image 2.1 on an RTX 5070 Ti 16 GB: 33s per 1024² image, down to 8. @ashxhart TensorFold's NVFP4 kernels, dropped into my normal ComfyUI as a loader node. Same workflow, same text encoder, sampler and VAE. 1024²: 33.0s to 8.2s (4.0x) 992×1216: 35.7s to 9.4s (3.8x) 480×608: 16.1s to 1.9s (8.4x) The catch: it's not pixel-identical. Same subject and clean images, sign text and hands are still right, but fine details drift slightly. To be fair, I have used this to generate hundreds of images and had a similar result even before the optimisation pass. FP8 stays closer to the original for strict editing and still gets a 2x speed boost. Side by side below. Judge for yourself. The recipe will be in the comments soon. DW, I am still cooking something for the DGX Sparks at the same time.
6
1
28
4,073