Abliterated Qwen3.8-27B picks an answer on sensitive prompts, then keeps arguing with itself until it hits the token cap.
ThinkFix is that abliteration plus one output row edit so it closes thinking when it's ready.
40 held out prompts, 8k cap, same weights: clean finish 22 → 33 hit the limit 17 → 5 75 min → 61 min median thinking ~3,000 → ~1,700 tokens
MATH / IFEval about flat. MTP kept, 56 tok/s vs 39 on a 3090. Q4_K_M is 16.8 GB.
huggingface.co/BoldingBuilds…
Abliterated Qwen3.8-27B picks an answer on sensitive prompts, then keeps arguing with itself until it hits the token cap.
ThinkFix is that abliteration plus one output row edit so it closes thinking when it's ready.
40 held out prompts, 8k cap, same weights: clean finish 22 → 33 hit the limit 17 → 5 75 min → 61 min median thinking ~3,000 → ~1,700 tokens
MATH / IFEval about flat. MTP kept, 56 tok/s vs 39 on a 3090. Q4_K_M is 16.8 GB.
huggingface.co/BoldingBuilds…
I've been trying to fix something frustrating about Qwen reasoning models: sometimes they think right up to the output limit and never write an answer. Not a refusal, not a wrong answer. Just nothing.
I tried fixing it by editing one row of the output layer, the row that scores the "stop thinking" token. Then I put the edited row back into a normal GGUF file.
On 200 math problems with a 4,096 token limit:
• Qwen3.8 27B: 31 blank answers → 0
• Qwen3.8 Flash-Next: 29 → 0
For comparison, llama.cpp's reasoning-budget flag also reached zero and scored about the same in these tests. The difference is that the edit lives in the file, so it works in apps that don't expose a reasoning-budget setting.
There are limits, and I tried to lay them all out: a few stray tags, one seed, tested mostly on llama.cpp.
Results, downloads, and the caveats:
huggingface.co/spaces/Boldin…
Built on @Alibaba_Qwen's models and @UnslothAI's GGUF conversions. If you've worked with local reasoning models, I'd like to hear what I got wrong.
Getting ready for tomorrow’s OpenAI DevDay! 🚀
Met some awesome new people today who, just like me, love to build, experiment with AI, and have fun doing it. Great conversations, exchanging ideas, and hanging out at Fort Mason right now.
San Francisco has been amazing so far and more new connections tonight!
@OpenAI@OpenAIDevs
Can’t wait for tomorrow. 🔥
@BoldingBuilds@_onmax@ThatGuySam
2026 is fucking wild.
Sitting on a plane to SF for OpenAl DevDay, using the in-flight WiFi to talk to my agent, which is remotely controlling my 4x3090 rig back in Texas + a RunPod GPU to run Qwen experiments while I'm in the air.
I love this shit.
The fix is one row of the output layer, the </think> token. 1,250 bytes.
I fit it on stock Bonsai's own reasoning so closing becomes likely once the answer is written and stays unlikely mid-thought.
The rest of v2 is the same refusal edit as v1, just stronger.
Also:
- 0% refusals, 0% over-refusal on safe prompts
- math, code, IFEval: no significant difference vs stock
- MMLU: -0.56 pts, smallest hit of the strong builds I tested (others -0.64 to -2.68)
- MTP file: 69 to 97 tok/s on a 3090
- tight budget? use reasoning_effort=medium
When Hikari, Heretic and dealignai do answer, their answers score about as well as v2's (0.96 to 0.99). The difference is how often they answer.
Overall score: v2 0.941, next best 0.738.
The chart going around from my Bonsai 2 report was thinking OFF.
With thinking ON, every ABLITERATED Bonsai 2 build in it has the same problem: it writes the answer inside its reasoning and never hands it over. You often get a blank reply.
Stock Bonsai doesn't (0 of 100 blank). v2 fixes it.
I gave each build the same 150 hard requests (default settings, 16k tokens). How many times it gave NO answer at all:
v2: 7
Blackfrost: 12
my v1: 23
Hikari: 34
Heretic: 41
dealignai: 48
Lower is better.
Answer quality: v2 0.941, next best 0.738.