dad, creator of LocalAI(localai.io) and Kairos (kairos.io) , ex @SUSE/@Rancher, ex-Gentoo Dev. vllm.cpp and APEX quants

Italy
a week ago I shared what I've been working on, vllm.cpp, vLLM's serving stack in C++ with no python in it. I got really a lot of feedback, and it seems I'm not the only one who wanted this to exist. We reached 800+ commits, 280 stars in just so little time. but where is it going? I got a lot of requests for what to support and what not to, so I took some time to write down why vllm.cpp exists, what it will be, and how you can contribute and it's already not just me. contributors showed up with hardware I don't own and did what I could not do because I don't have the hardware: - @DevAutomata got the @AMD backend going - @anothervariable got gemma-4-26B already running on two R9700s - @lu_zero_ did the whole Tenstorrent bring-up (!) - @jichiep has been looking at the mamba kernels - @filipsajdak got it building on Jetson Thor We are building a high-performance inference stack that runs everywhere and that you can embed in your own software. If you have hardware I do not have, that is the single most useful thing you can bring. If you want to know where are we going, how you can contribute, the vllm.cpp manifesto is in the link in the replies
9
12
56
3,697
parakeet-redux is a really interesting beast. You can extract the VAD head and then, you discover that has better recall than Silero. Despite the footprint compared to it, offline might have other uses. @vikhyatk cooked hard here, congrats for the release, and thanks for working in the open! By the way, all Parakeet-redux models are now in parakeet.cpp (links below). And you can use parakeet now as standalone VAD with both.
3
4
17
1,526
Your local agent can now hear what was said, who said it, and what happened around them. parakeet.cpp transcribes audio, separates speakers, detects 527 sound classes, and now enrolls and remembers named voices across recordings. One pass, local C++.
7
10
71
3,054
Enroll voices, and run the complete pipeline
3
265
Nemotron-3-Diarization was just released by @NVIDIAAI , and it runs in parakeet.cpp now, next to Parakeet speech recognition. It runs already on Apple with Metal, CUDA, @AMD and with anything that Vulkan supports. We (@LocalAI_API) also added sound detection in: @Xiaomi's CED models (527 sound classes), ported to C++ sitting alongside it. You can run it with a couple of click via API, all Local and open source. Parakeet.cpp now runs all three in one pass and you get the transcript, who is speaking, and what else you can hear. This scene from Sprite Fright took 2 seconds on a CPU.
3
4
18
833
In LocalAI it is one line from the model gallery then POST a wav to /v1/audio/diarization with include_text=true and you get speaker-attributed text. For sound tags: local-ai run parakeet-cpp-ced-base and /v1/audio/classification. There is a parakeet-cpp-realtime-scene entry too, which puts speakers and sounds on the realtime API next to the live transcript.
1
2
206
Speed check for the diarization part: 12 minutes, 3 speakers, in 5 seconds on a CPU. That is 2.9x NeMo's default on the same chip (1.5x against its compiled attention), and 99.9% of speech frames match its output. On a GPU the two tie at full precision.
1
3
124
Ettore Di Giacinto retweeted
I added full-body-control to my simulation using gem-x.cpp and GEAR SONIC (motion-bricks.cpp). The physically simulated character is controlled by my body movements and SONIC which tries to balance it. Inference is done on @LocalAI_API which takes in video frames and produces reference movements. Note that I don't think I'm getting the best performance here. There are probably some issues to iron between gem-x.cpp and motion-bricks.cpp.
2
6
31
1,659
Ettore Di Giacinto retweeted
Excited to support @mudler_it through the our Builder Program seeing vllm.cpp bring fast, local inference, LoRA support, and seamless integration with agent workflows is exactly the kind of open-source innovation we want to help accelerate. builders need flexible infrastructure they can actually rely on — whether they’re experimenting locally or scaling AI-powered products. Great work to the whole team! 🚀
A small update on what's going on in vllm.cpp: - Minimax and LTX 2.5 got Lora support. Now you can Load fused-model, or either make lora load in runtime with prompts in vllm.cpp - Support for Kev (@jaredpalmer), Laya and GLiNER2.5 (@fastinoAI ). Same SystemOne api, so you can switch quickly and have a fast Jev replacement, running locally - Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw runs at ~53 tok/s at high concurrency (up to c32) with DGX Spark I'm working now on the docs and making this a bit more easy to go, you'll also see all these models popping up in @LocalAI_API for a one-click install. A big shout out our friends at @regolo_ai , for sponsoring us as part of their Builder Program and giving access to their API, thank you 🫶! I've also added them to my agent harness, nib, so you can /login and use their API directly, links👇
1
1
4
410
A small update on what's going on in vllm.cpp: - Minimax and LTX 2.5 got Lora support. Now you can Load fused-model, or either make lora load in runtime with prompts in vllm.cpp - Support for Kev (@jaredpalmer), Laya and GLiNER2.5 (@fastinoAI ). Same SystemOne api, so you can switch quickly and have a fast Jev replacement, running locally - Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw runs at ~53 tok/s at high concurrency (up to c32) with DGX Spark I'm working now on the docs and making this a bit more easy to go, you'll also see all these models popping up in @LocalAI_API for a one-click install. A big shout out our friends at @regolo_ai , for sponsoring us as part of their Builder Program and giving access to their API, thank you 🫶! I've also added them to my agent harness, nib, so you can /login and use their API directly, links👇
3
4
17
1,661
Ettore Di Giacinto retweeted
I added kimodo.cpp animation gen & playback in latent3d.space. So you can describe a motion, then it'll generate it using a @LocalAI_API instance and play it back on the physically simulated character. Of course there are issues with finding animations that are physically possible, but that is part of the fun... maybe.
1
7
33
1,675
Ettore Di Giacinto retweeted
I now let my agent design its own control over his body - a tank :) Do I need a kill switch?
2
1
5
569