Really bullish on workload-aware inference! Database systems are of course a crown jewel of exploiting workload knowledge to achieve good performance. now in today's inference era, any system with a declarative interface and budding demand for LLM inference is probably a good candidate for workload-aware inference
Yep! A batch data processing pipeline usually has several LLM operations, and the API approach makes one call per row per operation, with each call served independently. If you know all the requests up front, you can plan the whole job, like reordering requests within an operation to share KV cache, running the most selective filter first, and reusing each document's KV cache across operations. Also beyond planning, you also want to change execution — eg if you know which requests in a batch share a prefix, you can use different attention kernels that read the shared KV from HBM once, instead of once per request