this is awesome!! video generated with code
1
54
This looks awesome!! thanks for such contributions
I’m back, with a little something… GPC 1, a new general-purpose classifier enabling: bounding boxes, poses, coordinates, direct mechanical control and so much more, all in milliseconds. Post-trained on millions of datapoints and thousands of unique problem classes learned directly from you guys posting on X, GPC-1 is a great tool for many things. Early user feedback is also good :) A new output class enables precise continuous range numerical outputs, so the model can output coordinates of bounding boxes, exact angles and distances, pose estimations, scores, and more. It goes a lot beyond typical classification which just pick a label. Excited to see where this goes. Available now!
1
80
Half if the people is still confused about JEV. So this is now my favorite way to consume AI news away from Hype 🤣
Holy crap. Rick and Morty just explained Jev AI to me better than any tech demo could.
98
Codacus retweeted
Qwen3.8-27B spent 21 days trying to build its own CUDA inference engine on a single RTX 3090. Mostly unsupervised. No human-written CUDA. Only around 12 human nudges across the entire experiment. The model wasn’t just given one giant prompt and left alone either. DeepSeek Harness managed the whole operation: Subagents. Roles. Handoffs. Context management. Compaction. HyperQwen served Qwen3.8-27B through vLLM while the agents worked on the inference engine. And the project accumulated some serious numbers over those three weeks: 180 subagents 230M tokens processed 1.7B cache-read tokens 699 context compactions 83 hours spent on context compaction The final engine didn’t beat llama.cpp. That’s important. HyperQwen managed roughly 250 prefill tok/s compared with around 700 tok/s for llama.cpp in the reported comparison. So this isn’t a story about an AI suddenly inventing a better CUDA runtime. It’s a story about how far you can push a local 27B model when you give it an engineering environment, persistent state and enough time. The funniest part was what happened when the agent started testing its own work. The RTX 3090 has 24GB of VRAM, and the same GPU needed to run both Qwen3.8-27B through vLLM and the CUDA engine being tested. That meant the workflow had to repeatedly: Stop vLLM. Run the benchmark. Restart vLLM. Continue the project. Then one subagent ignored the handoff protocol and killed vLLM while the orchestrator was still running. It basically shut down its own brain. A genuine AI suicide loop. And yet the larger system recovered and continued the project. That part is more interesting to me than whether the resulting inference engine beat llama.cpp. A quantized Qwen3.8-27B model maintained a real software engineering project for three weeks on consumer hardware. It spawned 180 subagents. Processed 230M tokens. Compacted its own context 699 times. And required only about a dozen human interventions. That’s a very different kind of benchmark. We’re starting to move from: “Can the model write code?” to: “Can the model maintain a software project long enough to actually get somewhere?” The answer here wasn’t “it built a better inference engine.” It didn’t. But it did demonstrate that a local Qwen3.8-27B can participate in a surprisingly long-running engineering workflow without a human sitting there writing the implementation. One clarification because the attached video can be misleading: The video is NOT the CUDA engine produced during the 21-day experiment. It compares HyperQwen, the inference backend used to run Qwen3.8-27B during the experiment, against stock vLLM on the same RTX 3090 using the same prompts. The 21-day experiment and the video comparison are separate things. The engine lost the speed race. The agent endurance experiment is still pretty wild.
Jurly
13
14
183
25,263
Codacus retweeted
"Can You Run Any LLM in Jev Mode Using llama.cpp?" @thecodacus "Jev answers in milliseconds, picks from your options instead of writing a paragraph... I added the same way of asking to llama.cpp, so it runs on the model already sitting on your drive..." piped.video/bcGO7xre46o
1
2
99
interesting idea
#ai anyone tried using connectome and fruit fly neurons mappings with a local LLM like qwen 3.8 27b or qwen 3.6 35B yet? @thecodacus
1
1
105
Codacus retweeted
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
640
1,572
15,547
5,444,109
Codacus retweeted
Fast meets open. 🚀 Qwen3.8-27B is now running on @cerebras with rapid inference. Try it now!
Qwen3.8-27B is now live at Cerebras speed. The dense, open-weight model from @Alibaba_Qwen scores 34 on the Artificial Analysis Intelligence Index—making it comparable to models such as GPT-5.6 Luna, DeepseekV4 Pro, and Claude Sonnet 4.6.
74
60
1,283
208,088
we have seen the video where Anthropic finds emotion directions inside Claude and nudges them to change its behaviour. Thought that was cool, so I tried it on Gemma 4 12B, little modification to the llama.cpp, and here we go
1
2
6
184
Codacus retweeted
"Is Frontier Class Local AI Finally Practical?" Qwen 3.8 Flash Next: "A 177-billion-parameter model, running on a five-year-old RTX 3060 with 12 GB of VRAM. Level with Claude Opus 4.8, going toe to toe with Opus 5 ..." @thecodacus @YouTube piped.video/IH8XmxiwliQ
1
1
83
Codacus retweeted
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: qwen.ai/blog?id=qwen3.8-flas… - Technical Report: github.com/QwenLM/Qwen3.8-Fl… - Hugging Face: huggingface.co/Qwen/Qwen3.8-… - ModelScope: modelscope.cn/models/Qwen/Qw…
461
1,039
7,962
1,928,554
Finally.. hoping its a sub 100b parameter model 🤞
Qwen announces Qwen3.8-Flash-Next, a new open-weight multimodal MoE model. 💜 The model will be released tomorrow and we are working on @UnslothAI day zero support. Qwen3.8-Flash-Next is built on the new Qwen4 architecture.
2
7
207
Watching this really breaks my heart. Years of documented knowledge, history, and civilization destroyed. are these the "good guys" who should only have the power to hold all these knowledge and the Frontier AI?
WTF AI companies are purchasing large quantities of used and rare books. Scanning their contents to train models. Then turning the originals to pulp. The sum of human thought, digitized and shredded. Every book that gets scanned disappears from the physical world permanently. The knowledge survives only as training data inside a system no one can hold, browse, or resell. What used to sit on a shelf for centuries now exists as weights in a model that might get deprecated next quarter.
2
3
7
200
Codacus retweeted
Qwen3.8 is launching and going open-weight soon!🌐 With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5. You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out. Can't wait to hear what you build. Stay tuned! 🚀  Token Plan international:qwencloud.com/pricing/token-… China:platform.qianwenai.com/prici…
1,317
3,284
24,180
8,427,639
Codacus retweeted
How can this man have so much aura?
空空
83
73
1,608
150,521
Youtube has some good moments, it gave me an interesting video suggestion, @thecodacus ran Qwen3.6 35b on an old 1060 and he got 17 tps. I guess I might start digging up some of my old hardware and see what it can do. That 1060 never dreamt in its entire life that it would prove to be so useful. Here's the link: piped.video/watch?v=8F_5pd…
1
2
72
I'm glad to be able to run a local assistant using Llama-cpp, Qwen3-coder-30B-A3B and Morefine RTX 4090M 16GB VRAM + 64GB RAM. > llama-server -m models/Qwen3-Coder-30B-A3B-Instruct-Q8_0.gguf --n-gpu-layers 999 --n-cpu-moe 28 --no-mmap --webui-mcp-proxy Thanks @thecodacus
2
1
3
163
⚡Build your own MCP Server using AI (educational) I use @Trae_ai and the MCP template from @cole_medin to create my own MCP servers, which then can be used within @n8n_io workflows. piped.video/WYD_OoPeBaE
2
5
30
4,287
Codacus retweeted
Created this game demo over the weekend without writing a single line of code! Built via bolt (Claude Sonnet 3.7) and powered by R3f + three.js Try it yourself - dodge enemies with real-time physics. Here's how I created the game + a cloneable character controller project 1/7
106
253
2,908
396,119
🔥New @boltdotdiy Release 0.0.7 just came out! piped.video/watch?v=rNhptI0V…
1
3
5
547