I build things to find out how things work. Local AI on consumer GPUs. Measure, don't assume. Currently: what's actually inside quantized LLMs.

Texas
Abliterated Qwen3.8-27B picks an answer on sensitive prompts, then keeps arguing with itself until it hits the token cap. ThinkFix is that abliteration plus one output row edit so it closes thinking when it's ready. 40 held out prompts, 8k cap, same weights: clean finish 22 → 33 hit the limit 17 → 5 75 min → 61 min median thinking ~3,000 → ~1,700 tokens MATH / IFEval about flat. MTP kept, 56 tok/s vs 39 on a 3090. Q4_K_M is 16.8 GB. huggingface.co/BoldingBuilds…
2
1
65
Abliterated Qwen3.8-27B picks an answer on sensitive prompts, then keeps arguing with itself until it hits the token cap. ThinkFix is that abliteration plus one output row edit so it closes thinking when it's ready. 40 held out prompts, 8k cap, same weights: clean finish 22 → 33 hit the limit 17 → 5 75 min → 61 min median thinking ~3,000 → ~1,700 tokens MATH / IFEval about flat. MTP kept, 56 tok/s vs 39 on a 3090. Q4_K_M is 16.8 GB. huggingface.co/BoldingBuilds…
2
1
65
I havnt played with Jev as much as I should have. I wonder if it could be used to optimize setting reasoning level for each prompt 🤔
50
Been a cool experience, even though the releases were a bit underwhelming
9
327
Pretty cool gift for OpenAI Devday
1
5
432
Devday feels like it’s more for business/consumer vs devs this year
30
I've been trying to fix something frustrating about Qwen reasoning models: sometimes they think right up to the output limit and never write an answer. Not a refusal, not a wrong answer. Just nothing. I tried fixing it by editing one row of the output layer, the row that scores the "stop thinking" token. Then I put the edited row back into a normal GGUF file. On 200 math problems with a 4,096 token limit: • Qwen3.8 27B: 31 blank answers → 0 • Qwen3.8 Flash-Next: 29 → 0 For comparison, llama.cpp's reasoning-budget flag also reached zero and scored about the same in these tests. The difference is that the edit lives in the file, so it works in apps that don't expose a reasoning-budget setting. There are limits, and I tried to lay them all out: a few stray tags, one seed, tested mostly on llama.cpp. Results, downloads, and the caveats: huggingface.co/spaces/Boldin… Built on @Alibaba_Qwen's models and @UnslothAI's GGUF conversions. If you've worked with local reasoning models, I'd like to hear what I got wrong.
1
2
402
I'll be at DevDay tomorrow. If you've run into runaway thinking with local models, I'd love to hear what you've tried.
1
45
It’s been awesome meeting everyone today who is as passionate about ai as I am! Looking forward to dev day tomorrow!
1
4
169
Josh Bolding retweeted
Getting ready for tomorrow’s OpenAI DevDay! 🚀 Met some awesome new people today who, just like me, love to build, experiment with AI, and have fun doing it. Great conversations, exchanging ideas, and hanging out at Fort Mason right now. San Francisco has been amazing so far and more new connections tonight! @OpenAI @OpenAIDevs Can’t wait for tomorrow. 🔥 @BoldingBuilds @_onmax @ThatGuySam
1
3
5
115
We don’t have these in East Texas 😂 First Waymo ride
55
Good morning San Fran
1
44
2026 is fucking wild. Sitting on a plane to SF for OpenAl DevDay, using the in-flight WiFi to talk to my agent, which is remotely controlling my 4x3090 rig back in Texas + a RunPod GPU to run Qwen experiments while I'm in the air. I love this shit.
2
3
101
Testing a one row Qwen3.8 modification that took blank outputs from 71/250 to 1/250.
1
30
Working on something that I think will be received well regarding 3.8 27b! Should have final results by morning! Stoked!
1
50
The fix is one row of the output layer, the </think> token. 1,250 bytes. I fit it on stock Bonsai's own reasoning so closing becomes likely once the answer is written and stays unlikely mid-thought. The rest of v2 is the same refusal edit as v1, just stronger.
1
66
Also: - 0% refusals, 0% over-refusal on safe prompts - math, code, IFEval: no significant difference vs stock - MMLU: -0.56 pts, smallest hit of the strong builds I tested (others -0.64 to -2.68) - MTP file: 69 to 97 tok/s on a 3090 - tight budget? use reasoning_effort=medium
1
62
Thanks @Hikari_07_jp for sharing the report this week. Their build still looks great on the thinking-off chart. This is a different test. Original report: huggingface.co/spaces/Boldin… v2, with every table: huggingface.co/BoldingBuilds…
1
140
When Hikari, Heretic and dealignai do answer, their answers score about as well as v2's (0.96 to 0.99). The difference is how often they answer. Overall score: v2 0.941, next best 0.738.
1
68
The chart going around from my Bonsai 2 report was thinking OFF. With thinking ON, every ABLITERATED Bonsai 2 build in it has the same problem: it writes the answer inside its reasoning and never hands it over. You often get a blank reply. Stock Bonsai doesn't (0 of 100 blank). v2 fixes it.
3
1
6
5,164
I gave each build the same 150 hard requests (default settings, 16k tokens). How many times it gave NO answer at all: v2: 7 Blackfrost: 12 my v1: 23 Hikari: 34 Heretic: 41 dealignai: 48 Lower is better. Answer quality: v2 0.941, next best 0.738.
1
123