Agentic coding tasks burn ~1,000x more tokens than chat, and more tokens don't always buy better outcomes. What does? The harness around the model. New explainer, from Hermes and Muse to Claude Code and omp.
When you start chatting with a model or using an agent, everything you share and that is generated back to you is initially stored in the KV cache.
But what is a KV cache? How does it work and why does its size and precision matter?
Check out my latest.
Have you wondered how AI works? Most of today's AI is powered by what are called large language models. This guide breaks down how they work in an approachable and easy to understand way.
Running a 70-billion-parameter language model on a single graphics card sounds impossible. A model that size needs roughly 140 gigabytes of memory in its native 16-bit format, far beyond the 24 gigabytes a consumer GPU offers. But EXL3 made it possible!
Nvidia Blackwell GPUs introduced native hardware support for 4-bit floating-point computation through a format called NVFP4. The results are concrete: 2-3x higher throughput than FP8, 3.5x less memory than BF16, and accuracy within 1-2% on large models.