Hard Work and Guts

Austin, TX
Pinned Tweet
2 months ago, AI was a chatbot that I occasionally used for random questions. Then I came across Ahmad’s posts about how it works and what it can do, and I knew immediately this was something I wanted to dedicate all my time and resource into and finally, my own sovereign AI. I started out with the goal of having my own AI assistant that is 100% private, that I can access anywhere using text, voice and vision. Currently building out the second brain infrastructure that can continuously improve itself, with actual long-term memory so it doesn’t just end up being another glorified chatbot. Afterwards, tie it into genomics research and VCF automation, and design my private console to drive and monitor the entire infrastructure. I’ll start sharing my journey in a way that isn’t just technical jargon, my hope is that it’ll inspire more people to get into AI to improve their everyday lives.
92
35
569
83,851
So this monstrosity of a RTX 4500 has been driving me nuts, its passive cooled and didn’t want to use those high static fans with 15,000rpm. Tried custom 120mm shround but it barely keeps it below thermal throttle, so took it apart to measure it to make a waterblock for it Turns out it’s only half full, so I think I can just cut it open or in half, slap a threadripper waterblock on top of the heatsink and cool it that way. Will let yall know if it works or not when Optimus block arrives.
8
1
49
2,847
For those with older motherboards that has Gen 2/3/4 PCIe slots, you can use latest GPU without any noticeable impact on performance That is until TP comes in, and even then, decode will remain barely changed. Each sync is tiny (about 10 KB), so cost is waiting x per sync (latency-bound), not bandwidth. Slower links add few microseconds per sync, repeat 128 times per word. At Gen2, TP2 adds 0.50 ms per word, TP4 adds 1.41 ms which is negligible Where it will hurt is prefill. Reading prompt (prefill) processes 4,096 tokens in this example, so each sync carries roughly 42 MB instead of 10 KB. That's about 5 GB crossing the slot per 4K tokens of prompt, roughly 57 GB for a 44K prompt on TP2. That is pure bandwidth problem, halve the link speed and the waiting roughly doubles. So for single GPU, feel free to use any Gen PCIe slots you have, and avoid any split card on it!
4
6
40
1,747
Just got my LLC approved here in Texas, initially planned to do AI genomics work with wife but thinking about expanding to other areas like AI measurement/research. Always liked to measure things for myself than relying on others, too often you find public numbers are rather deceptive or inaccurate Now that labs settled down somewhat, I’ll be going back to publishing model analysis and difference between quants. Aside from 6000s, I have 4500,5090,5080,5070,3090,1080ti,980ti that can target every VRAM rungs so I can see what fits with how much context/KV. With TP8, I should be able to do these things at scale so if there are any models you want me to look into, always feel free to reach out. My work here will always be free, transparent, and very random.
16
3
95
2,608
Some of yall asked what QDC or quick disconnect couplings were so here it is. This right here is what makes watercooling infinitely easier, allows you to disconnect components without having to drain, and without a single drop of leak. Crucial to use.
3
57
2,778
Finished the waterloop on the 2nd server, couldn’t keep thinking about that skateboard meme where dude does 10 steps to land on his face. This is the definition of unnecessary, even I was like wtf am I doing, I could do this with two tubes and pump/rad on the ground but that would be make too much sense for me. Anyways, all of this gets taken apart when rest of TP8 parts come so there’s that too
41
7
203
9,375
My OCD is kicking in hard looking at all the imperfections but I had to step away and call it finished since both hosts will be dismantled in few weeks.
1
6
608
Root/leafs design for my TP8 build, cross GPU traffic never hitting CPU bottleneck. $9k in components (cables alone was $1300) and additional $60k for 4 6000s Full sovereignty, privacy, run some of the biggest models that answers to no one but myself, thats worth something Total will probably come out to $150,000 for TP8 watercooled host with 9985wx, 768 DDR5 6400, 768 VRAM, 20TB+ U2 NVME, 100Gb NIC when all said and done. I still would like couple of H200s someday
12
1
60
2,846
Giving up on H200 and going RTX 6000 TP=8 using 3 of these switches. Managed to pry the last 3 so I can do root/leaf architecture This is going to be an enormous undertaking, powering 4 additional 600w GPUs and creating watercooling loops is going to require significant effort
19
2
137
6,933
If you are getting into local AI, I just have one recommendation. Watercool your GPUs Inference is one of the most taxing thing you can put any GPUs through, even more so if you are doing any sort of training that runs for hours on end. This leads one or combination of these things. GPU temp hitting or hovering near thermal throttle Loud noises Torturing your precious GPUs, shortening life Lower performance All of those can be mitigated to some degree by various means, but only one solution solves all of it which is Watercooling. I’ll go into detail on how for every one, for this main post all I can say is it will, I have 8 watercooled GPUs next to me running max overclock that doesn’t go above 65C no matter how much I abuse it Now the main objection will be, ‘I don’t want to risk my $$$ GPUs on something so dangerous’ and you would be correct, if you weren’t taking certain steps that can make it absolutely safe in mechanical term. The biggest risk you can eliminate is, don’t make it look good. This is by far the biggest killer of watercooled setups, since it opens up way too many areas of failure on or around your GPUs. Trust me, I nearly lost four 6000s while soaked in leaked coolant running drafter training for 5 hours without knowing because a hard tube cracked right above them (all survived, made a post awhile back, modern GPUs are extremely durable) Make it look ugly, use quick connect fittings and just connect some tubes together away from your GPUs and you’ve now eliminated any chance of leaks damaging anything. Don’t be afraid to open GPUs up, you can touch them all day and it will not break. Put a waterblock on it, connect to a pump, reservoir and radiator and now you can get 15-20% decode while never seeing high temp ever again no matter what you do I’ll make a post later on some areas and anything else yall want to know about so you have the full picture
44
17
370
28,232
x.lingyaoai.com/net_termina/status/210… Forgot to include this, here’s what you actually get out of it
Quick follow up to watercool post, to show why it’s worth doing #1 Noise. With noctuas on radiators, you won’t notice it’s on #2 Temp, it will refuse to rise higher than 60C unless you burn at 600w for few hours, then maybe 65C which allows you to - #3 Overclock, permanently
5
478
Quick follow up to watercool post, to show why it’s worth doing #1 Noise. With noctuas on radiators, you won’t notice it’s on #2 Temp, it will refuse to rise higher than 60C unless you burn at 600w for few hours, then maybe 65C which allows you to - #3 Overclock, permanently
6
3
42
2,725
Hard to believe it’s only been 3 years, what’s it going to look like 3 years from now?
2
4
39
1,886
I saw a comment recently that said ‘harness is everything’, and it couldn’t have been more accurate when it comes to models with multiple thinking modes. Using Bonsai 2 in my example since it encapsulates this point so well. Spread on coding bench scores are simply out of this world depending on so many factors such as your app, harness, thinking mode, token cap, all of which I’ve tested to get to the bottom of this. Bonsai 2 does code at 27B level, on medium with minimum 8k+ cap which is not the default setting for most app/harness that was hurting performance badly. 27B have similar issues to be fair. Generally xhigh only performs best when given outrageously high cap which nobody will sit through, so if you are getting bad results with thinking mode on, try higher cap and different modes to see if it helps your use cases.
5
3
49
2,264
For final 1080ti experiment, I wanted to do the card justice so held nothing back Bonsai 2, custom kernel + native MTP drafter trained on hidden states. Pascal chads can now enjoy this tiny, extremely capable model 86 code, 98 math, 75 prose tok/s 32K/64K context 1 GPU, 128K+ w/ 2 (No NV Link needed) Please use medium thinking mode with 8k+ budget cap for true 27B level experience, I’ll make a separate post about this tonight. Lastly, I’m giving away this 1080ti no strings attached. Need to get back to modern GPU experiments so if you’d like one, let me know and I’ll have Claude randomly pick someone.
33
10
107
4,719
Can you tell me your specific use cases and how it fails on your model? Doing research on model capabilities, trying to determine how much of your failure is harness, serving config, or straight up model being not good enough. If you reply with any much detail as you are willing to share and on which model, I’ll test every single one myself and let you know if model/config is your issue.
8
1
16
1,341
GPU bandwidth tier list based on how many kidneys you need to sell to buy one and how much bandwidth you get per kidney. (both AMD and NVIDIA cards this time) Included sweeps/s as a bonus. You can replace KU with anything btw, 1oz of gold, $5,000 cash, etc and speed value is still accurate.
7
2
65
5,357
About sweeps/s, it measures one thing and one thing only. The speed at capacity for any given GPU running a dense model. I gave an extremely example last time, but here’s more useful one. 5090 vs 6000 both have same bandwidth, but load both up that takes 80% VRAM, 5090 will run it much faster since it has some of the highest sweeps/s out of all GPUs. Yes model size will be different, its just a speed measurement at capacity beyond simple bandwidth.
4
518
1080ti, the emperor of GPUs, time to rise again 237 tok/s short — 160 tok/s long humaneval Full 128k context Everything you want to know in recipe below Had to sacrifice a GPU in the process and many sleepless nights but I’m a man of my word.
20
10
119
9,092
Forgot to post prefill and ladder rung bench.
7
497
Also spent too many hours building kernel/drafter for Bonsai 2, but decided not to share it here. Will do another post dedicated to it, as it’s more sensitive topic. However, if you’ve had bad experience with it, run it on MEDIUM thinking mode ONLY and with minimum 8k budget or don’t use it at all
3
1
8
913