A tower defense game, built by two different models on the same coding-agent task: one run cost $1.50, the other 27 cents.
@bgchun, founder and CEO of FriendliAI, walks through that gap, and what closing it takes on the inference side, in "The Frontier AI Inference Cloud for Agents," on
@aiDotEngineer's YouTube. The talk lays out why serving agents is a different problem from serving chat, and how FriendliAI rebuilt its stack around that difference.
- Two trends colliding. Open-weight models have reached frontier-level quality right as agents move into production at scale, which is what makes the economics below possible.
- The cost gap, made concrete. On the same task, building a tower defense game with a coding agent, GLM 5.2 on FriendliAI cost 27 cents against $1.50 for Opus 4.8, a comparable result at a fraction of the price.
- Agents aren't chat with more requests. A session runs as a loop of plan, act, observe, with tool calls interleaved between model calls, and consecutive steps sharing a huge prefix of context that's expensive to recompute.
- Four pillars of the rebuilt stack. Prefix caching, KV cache management, cache-aware routing, and agent-aware optimization, all aimed at end-to-end task latency rather than the latency of any single request.
- Continuous batching and Orca. Gon's research team invented continuous batching, now standard across the industry, and its Orca work helped shape vLLM.
- A number from production. Kilo Code, testing FriendliAI against other providers on GLM traffic, found it consistently seven times faster with a lower error rate.
- Three ways to consume it. A serverless model API, dedicated endpoints with SLAs, or running the same stack on your own GPUs.
I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!