Control plane for agents & engineers to provision compute and run training & inference across NVIDIA, AMD, and other chips — on clouds, Kubernetes, and on-prem.

github.com/dstackai/dstack
Pinned Tweet
Inference serving is increasingly open source. But inference optimization still happens inside each provider, and the optimized deployment remains tied to its proprietary stack. This has to change. Introducing Presets: an open-source toolkit for agent-based inference optimization and a portable format for the result. Deploying optimized Kimi K3 to any cloud, Kubernetes cluster, or bare-metal fleet, whether on NVIDIA, AMD, or other silicon, should be as simple as deploying a Docker image. dstack.ai/blog/presets/
1
5
12
2,115
dstack retweeted
dstack 0.22.2 is out! You can now provision compute on @daytonaio and orchestrate training and inference with NVIDIA B300/H200 and AMD MI355X GPUs. Presets now use cache-aware routing even without PD disaggregation. Release notes: github.com/dstackai/dstack/r…
1
4
11
3,566
dstack retweeted
We’ll be in San Jose for @PyTorch Conference North America Oct 20 - 21. 🎤 Talks on Triton, MoE fine-tuning + model customization 🤖 Booth demos + challenges 🍻 AI Infra Night with @dstackai, @SGLang (@lmsysorg) Register: 🔗 luma.com/6j69xyal #PyTorchCon ​​#PyTorchFoundation
4
4
10
3,950
Spin Daytona GPU sandboxes directly from @dstackai
You can now provision Daytona GPU sandboxes directly from @dstackai. On-demand + spot, per-second billing, B300s included.
2
6
240
dstack retweeted
Going to @PyTorch Conference in San Jose next month? @CrusoeAI and we are hosting an evening on day one, October 20, 6:30 PM, right after the sessions. We're gathering GPU experts, AI researchers, AI infra engineers, and anyone into AI infra and heterogeneous AI compute. Space is limited: luma.com/6j69xyal
1
4
6
3,067
dstack retweeted
dstack 0.21.3 is out! Presets: optimized inference configurations can now be pushed to a private registry and later deployed to any cloud, Kubernetes, or on-prem cluster. Backends: @seeweblive GPU cloud is now supported. Plus improvements to services and gateways. Release notes: github.com/dstackai/dstack/r…
2
4
259
dstack retweeted
Here's one example of using the presets toolkit to optimize Qwen3.8-27B on a single @AMD MI300X. Through compound learning across linked sessions and source-level patches, performance improved from 311 to 495 tok/s: +59%. The resulting preset can be deployed on any AMD cloud, Kubernetes cluster, or bare-metal fleet.
Inference serving is increasingly open source. But inference optimization still happens inside each provider, and the optimized deployment remains tied to its proprietary stack. This has to change. Introducing Presets: an open-source toolkit for agent-based inference optimization and a portable format for the result. Deploying optimized Kimi K3 to any cloud, Kubernetes cluster, or bare-metal fleet, whether on NVIDIA, AMD, or other silicon, should be as simple as deploying a Docker image. dstack.ai/blog/presets/
3
3
402
dstack retweeted
Inference serving is increasingly open source. But inference optimization still happens inside each provider, and the optimized deployment remains tied to its proprietary stack. This has to change. Introducing Presets: an open-source toolkit for agent-based inference optimization and a portable format for the result. Deploying optimized Kimi K3 to any cloud, Kubernetes cluster, or bare-metal fleet, whether on NVIDIA, AMD, or other silicon, should be as simple as deploying a Docker image. dstack.ai/blog/presets/
1
5
12
2,115
Inference serving is increasingly open source. But inference optimization still happens inside each provider, and the optimized deployment remains tied to its proprietary stack. This has to change. Introducing Presets: an open-source toolkit for agent-based inference optimization and a portable format for the result. Deploying optimized Kimi K3 to any cloud, Kubernetes cluster, or bare-metal fleet, whether on NVIDIA, AMD, or other silicon, should be as simple as deploying a Docker image. dstack.ai/blog/presets/
1
5
12
2,115
Here's one example of using the presets toolkit to optimize Qwen3.8-27B on a single @AMD MI300X. Through compound learning across linked sessions and source-level patches, performance improved from 311 to 495 tok/s: +59%. The resulting preset can be deployed on any AMD cloud, Kubernetes cluster, or bare-metal fleet.
1
63
dstack retweeted
i've been saying for a while now that we need a good harness for tuning inference, but also deploying it. @dstackai has picked up the ball and done it. glad to partner with them for years now. keep up the great work.
Inference serving is increasingly open source. But inference optimization still happens inside each provider, and the optimized deployment remains tied to its proprietary stack. This has to change. Introducing Presets: an open-source toolkit for agent-based inference optimization and a portable format for the result. Deploying optimized Kimi K3 to any cloud, Kubernetes cluster, or bare-metal fleet, whether on NVIDIA, AMD, or other silicon, should be as simple as deploying a Docker image. dstack.ai/blog/presets/
1
3
12
1,370
dstack 0.21.2 is out 🚀 Distributed tasks now support node groups: one run can mix roles and hardware. Each group defines its own nodes, resources, commands, and ports, e.g. a CPU head node driving GPU workers in a single task. Also in this release: the @AMD Developer Cloud VM image moves to ROCm 7.14, plus a major bugfix. Full release notes 👇 github.com/dstackai/dstack/r…
1
1
2
233
Node groups are built for workloads like RL training: training jobs, inference, and sandboxes can now run side by side in one run, each on its own hardware. Here's an example that uses groups to run @raydistributed on dstack, fine-tuning an agent with RAGEN and @verl_project: dstack.ai/docs/examples/trai…
1
2
2
145
dstack retweeted
dstack 0.21.1 is out! Presets: the inference optimization agent can now patch source code (serving framework, kernels, etc). Also, sessions can be linked, so each new one starts from the previous best and pushes further. Gateways: replicas are now fault tolerant and support in-place scaling. Plus lots of other improvements: github.com/dstackai/dstack/r…
1
3
3
1,707
Love it! Works on @tenstorrent. Go @dstackai
Super excited about this release. Most of the work went into presets. If you spend time optimizing inference, this is for you. You point the agent at a model and it figures out how to serve it best, down to patching kernels when flags aren't enough. A built preset can be deployed anywhere. Works on @nvidia, @AMD, and @tenstorrent, any cloud, K8S, or bare metal. dstack.ai/docs/concepts/pres… A longer write-up with results from our own sessions is coming soon.
2
9
1,266
Super excited about this release. Most of the work went into presets. If you spend time optimizing inference, this is for you. You point the agent at a model and it figures out how to serve it best, down to patching kernels when flags aren't enough. A built preset can be deployed anywhere. Works on @nvidia, @AMD, and @tenstorrent, any cloud, K8S, or bare metal. dstack.ai/docs/concepts/pres… A longer write-up with results from our own sessions is coming soon.
dstack 0.21.1 is out! Presets: the inference optimization agent can now patch source code (serving framework, kernels, etc). Also, sessions can be linked, so each new one starts from the previous best and pushes further. Gateways: replicas are now fault tolerant and support in-place scaling. Plus lots of other improvements: github.com/dstackai/dstack/r…
3
3
15
2,731