this week was really crazy in AI. And I have only a DGX box for this, not sure how long it will last.
What I'm doing in parallel for vllm.cpp:
- Qwen 3.8 27b benchmarks (I'd like to take some numbers to show you)
- Qwen 3.8 2.4t fitting on a single Nvidia DGX box (yes, we will have it, even if slow)
-
@Lightricks LTX2.5 support
- Index-tts support
-
@MiniMax_AI Music 3
-
@NVIDIAAI 's new Nemotron release
- Dspark just landed
- dots3-preview support in vllm.cpp
And since I'm having issues in overbooking the box with multiple agents, I'm doing a resource-controller api in parallel to control this..