Inference serving is increasingly open source. But inference optimization still happens inside each provider, and the optimized deployment remains tied to its proprietary stack. This has to change.
Introducing Presets: an open-source toolkit for agent-based inference optimization and a portable format for the result.
Deploying optimized Kimi K3 to any cloud, Kubernetes cluster, or bare-metal fleet, whether on NVIDIA, AMD, or other silicon, should be as simple as deploying a Docker image.
dstack.ai/blog/presets/