Today I'm releasing LlamaStash 0.0.2: a zero-overhead, terminal-native launcher for llama.cpp.
One Rust binary that's a TUI, a CLI, a daemon, and an OpenAI-compatible proxy.
Demo below 🧵
LlamaStash v0.6.0 is out 🦙
Highlights:
- a request that doesn't fit now unloads idle models to make room, and 'presets save --idle-ttl 0 --preload' keeps a model warm.
- Claude Code's /effort now reaches llama.cpp models.
llamastash.dev
I benchmarked engines for Qwen 3.8-Flash-Next on Strix Halo (Flow Z13, 128 GB). Halogen with its native weights is fastest at 39-40 t/s decode, then gufo and CIRU. But Halogen is closed source. gufo is open source and loads 4x faster from cold.
teddit.net/comments/1wu0m53
LlamaStash v0.5.0 is out 🦙
Highlights:
- generic backend: any OpenAI-compatible server you declare in config.yaml becomes a managed model. I run Halogen and gufo this way now.
- both beat llama.cpp on Qwen3.8 Flash-Next.
llamastash.dev
LlamaStash v0.4.0 is out 🦙
Highlights:
- SGLang backend next to vLLM for safetensors repos. 'start owner/repo --backend sglang' picks it per launch.
- safer memory checks on unified-memory hosts, so a launch that would freeze the machine is refused.
llamastash.dev
Qwen 3.8 replaced Claude Opus for my coding. 27B and Flash Next on a 128GB Strix Halo laptop, no cloud subscription. Flash Next scores 40 on the AA index against 42 for Opus 4.8. Decode 10-15 tok/s, prefill is the real cost.
deepu.tech/local-ai-qwen3.8-…
LlamaStash v0.3.0 is out 🦙Highlights:- one model, several copies, each under its own name. 'start qwen3 --name coder', and it answers to qwen3@coder on the proxy, CLI and TUI.- 'llamastash run model.yml', a preset file you can commit.llamastash.dev
LlamaStash v0.2.0 is out 🦙
The big change: a vLLM backend. The safetensors repos in your HuggingFace cache were invisible before; now they show up in list, launch from the TUI, and answer on the proxy. llama.cpp still owns GGUF.
llamastash.dev
LlamaStash v0.1.0 is out 🦙
The big change: MTP speculative decoding, on by default. Roughly 2x faster decode on models that support it.
Also new: pick a CUDA/ROCm/Vulkan build per launch, and one JSON shape across list and show.
llamastash.dev
How much local LLM can you run on an AMD Strix Halo with 128GB memory?
I managed to fit DeepSeek v4 Flash 284B and Gemma 4 E2B on GPU, Whisper and Qwen3.5 4B on NPU.
#strixhalo#amd#deepseek#llamastash
LlamaStash v0.0.6 is out 🦙
A experimental ds4 backend runs @antirez DeepSeek-V4 GGUFs through DwarfStar (ds4)
Plus: Lemonade on by default, and saved presets that auto-apply.
llamastash.dev#ds4#AI#deepseek
LlamaStash v0.0.5 is out 🦙
New: named launch presets. Tune a model's launch knobs once, name them, reuse them, per-model or per-arch. They live in plain config.yaml, so you can hand-edit, comment, and commit them to your dotfiles.
llamastash.dev
LlamaStash v0.0.4 is out 🦙
- Auto launch is now the default: llama.cpp's --fit sizes context and GPU offload.
- A browser UI on a stable port
- Anthropic Messages API support.
llamastash.dev
KDash 2.0 is out. The Kubernetes terminal dashboard now does more than watch.
- Delete, edit, scale, restart, cordon, port-forward, all from the TUI
- Action menu with confirm prompts
- New themes + live switching
Built in Rust 🦀
github.com/kdash-rs/kdash#Rust#k8s#DevOps
LlamaStash is multi-backend now 🦙 (v0.0.3)
llama.cpp stays the zero-overhead default. An experimental, opt-in Lemonade backend unlocks the AMD NPU, vLLM, ONNX and more.
Plus vision/audio models, a LAN proxy with bearer auth, and multi-GPU support.
llamastash.dev
How much overhead does an LLM launcher add? I matched flags and measured across AMD APU, Apple Silicon, and NVIDIA. Wrapper overhead: within 1% of raw llama-server. TTFT is where Ollama and LM Studio diverge. Methodology + JSONs reproducible.
deepu.tech/benchmarking-llam…
Today I'm releasing LlamaStash 0.0.2: a zero-overhead, terminal-native launcher for llama.cpp.
One Rust binary that's a TUI, a CLI, a daemon, and an OpenAI-compatible proxy.
Demo below 🧵
Install:
curl -fsSL llamastash.dev/install.sh | sh
or
irm llamastash.dev/install.ps1 | iex
or
brew install llamastash/llamastash/llamastash
or
yay -S llamastash
or
cargo install llamastash
Then `llamastash init` and you're chatting locally in minutes.