A small update on what's going on in vllm.cpp:
- Minimax and LTX 2.5 got Lora support. Now you can Load fused-model, or either make lora load in runtime with prompts in vllm.cpp
- Support for Kev (
@jaredpalmer), Laya and GLiNER2.5 (
@fastinoAI ). Same SystemOne api, so you can switch quickly and have a fast Jev replacement, running locally
- Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw runs at ~53 tok/s at high concurrency (up to c32) with DGX Spark
I'm working now on the docs and making this a bit more easy to go, you'll also see all these models popping up in
@LocalAI_API for a one-click install.
A big shout out our friends at
@regolo_ai , for sponsoring us as part of their Builder Program and giving access to their API, thank you 🫶! I've also added them to my agent harness, nib, so you can /login and use their API directly, links👇