Passionate about markets and startups. Enjoyer of local AI, security fanatic and builder of things. Working hard to make this world a better place.

United States 馃嚭馃嚫
Agentic coding tasks burn ~1,000x more tokens than chat, and more tokens don't always buy better outcomes. What does? The harness around the model. New explainer, from Hermes and Muse to Claude Code and omp.
Article

AI Agents, Explained: The Model Is Only Half the Story

You have probably asked a chatbot to summarize a document or draft an email. Now the same kind of technology can open a code editor, find a bug, fix it, run the tests, and leave a note explaining what

31
6
35
9,469
When you start chatting with a model or using an agent, everything you share and that is generated back to you is initially stored in the KV cache. But what is a KV cache? How does it work and why does its size and precision matter? Check out my latest.
Article

The Notebook Behind the Model: A Beginner's Guide to the KV Cache

Every time you talk to a language model, it keeps a notebook. Everything you have said, and everything it has said back, gets an entry. That notebook is the KV cache. Once you understand it, a pile of

9
8
27
15,076
Running a 70-billion-parameter language model on a single graphics card sounds impossible. A model that size needs roughly 140 gigabytes of memory in its native 16-bit format, far beyond the 24 gigabytes a consumer GPU offers. But EXL3 made it possible!
Article

EXL3: How ExLlama's Trellis Quantization Shrinks Large Language Models

Running a 70-billion-parameter language model on a single graphics card sounds impossible. A model that size needs roughly 140 gigabytes of memory in its native 16-bit format, far beyond the 24

9
5
29
10,943
Nvidia Blackwell GPUs introduced native hardware support for 4-bit floating-point computation through a format called NVFP4. The results are concrete: 2-3x higher throughput than FP8, 3.5x less memory than BF16, and accuracy within 1-2% on large models.
Article

NVFP4 on Blackwell: What 4-Bit Floating Point Actually Delivers

NVIDIA's Blackwell GPUs introduced native hardware support for 4-bit floating-point computation through a format called NVFP4. The results are concrete: 2 to 3 times higher throughput than FP8, 3.5

10
10
70
47,592