Vibe Coding + Vibe Marketing + Vibe Selling, Startups.

I ran the Laya model on llama.cpp to test the game of Snake, and the results were excellent. The speed was fast, and even on my laptop with an AMD Radeon 780M Graphics, it ran smoothly. Thanks @chenchengpro game code. #Jev #Laya #TypeSafe #llamacpp #LocalLLM
1
4
296
fudingyu retweeted
We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions! Here is a breakdown of what’s being fixed and updated in this release: 🧵👇
130
322
3,313
720,802
fudingyu retweeted
WSL container is now available for public preview. That means you can create, run, test, and debug Linux containers directly through WSL on Windows. Here’s how to get started + a few ways to try it 🧵
33
291
1,177
98,629
fudingyu retweeted
Introducing InfiniteDiffusion, my independent paper accepted to #SIGGRAPH2026! I have one RTX 3090 Ti. No funding, advisors, or team. By day I'm a new grad SWE at Walmart. The paper has two main contributions: - InfiniteDiffusion: a new approach to infinite generation with diffusion models. - Terrain Diffusion: the world’s first learned procedural terrain generator. Here’s why this matters, and how they are connected. 🧵
152
626
6,287
944,838
Windowsの効果音を口だけで表現するアカペラグループが凄すぎる
78
1,429
7,842
846,444
fudingyu retweeted
Found this Korean guy using AI video to make ads for products that should exist. No idea what he’s saying, but I’m a fan 😅 (this is 404product on IG)
216
442
4,096
1,384,523
fudingyu retweeted
🚀We’re excited to officially release Hy-Memory — a powerful memory plugin built specifically for long-term collaborative Agents like OpenClaw. More than a retrieval tool, it becomes your Agent’s true “Second Brain.” Powered by a 6-layer memory framework × System1/System2 dual system × three-layer evolutionary chain, Hy-Memory lets Agents remember durably, accurately, lightly, and understand you better. ➡️Solves memory fragmentation ➡️70%+ fewer memories ➡️45%+ higher info density per memory ➡️35% less token usage on ultra-long contexts ➡️20% faster memory updates. Upgrade your Agent’s memory today! 📷Project & Download: memory.hunyuan.tencent.com/ 📷 OpenClaw Docs: memory.hunyuan.tencent.com/o…
41
80
934
60,770
fudingyu retweeted
The next evolution of Hermes Agent is here! Introducing Hermes Desktop: everything you love about Hermes, now native on your machine. First demoed in Jensen's GTC keynote, it's now in public preview.
1,220
1,432
12,674
5,876,106
fudingyu retweeted
Thats awesome! A developer used Al-powered 4D Gaussian Splatting to convert flat video footage into dynamic 3D spatial scenes. The system reconstructs different camera angles and depth information from ordinary footage, making it possible to navigate scenes in three dimensions. We're moving from recording video to digitally recreating reality itself.
29
182
939
73,763
fudingyu retweeted
I didn't expect DeepSeek v4 PRO (not Flash) to run well on the Mac Studio M3 Ultra with 512GB of RAM. This is 2 bit quantized with the same DwarfStar recipe used for Flash. 433GB GGUF file. 130 t/s prefill, 13 t/s generation. Prefill in the video is low because small prompt.
53
82
1,064
173,018
fudingyu retweeted
If we ever figure out how to load ONLY the active params of an MoE into the GPU instead of the full weights, it's game over. Data centers would see a 100x efficiency boost. And we could literally run 1T models like Kimi locally on just 32GB VRAM. Yeah I know it's basically impossible right now, but who knows what the future holds. Let me dream.
89
34
816
72,907
fudingyu retweeted
DS4 running on DGX Spark (GB10 / CUDA), private branch for now. 12 tokens/sec, the memory bandwidth is limited in this system, at 270GB/sec. But prefill is ways more alighed to M3 Max at ~200 t/s. I'll release when more mature, but it is almost sure that it will get merged.
48
72
774
84,469
fudingyu retweeted
Welcome to DS4, a specialized inference engine for DeepSeek v4 Flash. github.com/antirez/ds4 This project would have been impossible without the existence of llama.cpp and GGML and the work of @ggerganov and all the other contributors. Thanks!
47
218
1,494
202,422
My 4090 went from 26 -> 154 tok/s Qwen 3.6 27B🤯 Same GPU. Same Q4_K_M . No FP8, no extra quant. The unlock: ik_llama.cpp + speculative decoding using Qwen3-1.7B as the draft model. 85% acceptance rate. Full config + benchmarks 👇🏻
81
146
1,637
128,309
fudingyu retweeted
Kimi K2.6 + DFlash: 508 tok/s on 8x MI300X 5.6x throughput improvement over baseline autoregressive serving 90 tok/s → 508 tok/s on the same hardware, same model, zero quality loss huggingface.co/florianleiber…
12
33
348
36,603
fudingyu retweeted
Everyone says k2.6 is unusable. I had it build a pokemon style battle app. 1 prompt. Kimi cost 8 to 10 times less on tokens than opus 4.7. Which is better?
39
42
604
127,541
fudingyu retweeted
We're open-sourcing FlashKDA — our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. Achieves 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention. Explore on github: github.com/MoonshotAI/FlashK…
45
186
1,799
221,675