500M tokens a day in March 2026.
15B tokens a day yesterday.
That is daily volume across
@ionet hosted models on
@OpenRouter, served on an almost same mix of H100, H200 and B300 GPUs since Day 1.
Along the way, we became day zero launch partners for open source reasoning models including
@Zai_org,
@Kimi_Moonshot,
@Alibaba_Qwen , Nemotron,
@MiniMax_AI and GPT OSS.
Nvidia has built something useful with the recent Nemotron 3.5 Lightning. A fast, capable and efficient open model, with weights, datasets and recipes that give teams like ours room to optimise.
Yesterday, the OpenRouter provider token share showed
@ionet at 64% token share for Nemotron 3.5 Lightning, the largest provider share.
Effective prices per million tokens:
Input: $0.0335
Output: $0.179
Our volumes and price reflect both usage and affordability. We are serving this workload profitably on an inference stack built and operated end to end in house.
Our small team still runs two to three experiments every week. We track tool calling changes, frontier lab updates and what communities on vLLM and Sglang are discussing, then test what matters on our own workloads.
Demand changes with time zones and days of the week. Launch excitement fades. New providers arrive and price wars start. You get used to revisiting your assumptions.
With GPU rental costs rising, compute efficiency and reliable tool calling matter more. A low token price only helps if the task gets done correctly.
Open models are becoming capable enough to build around and economical enough to serve at scale. There is a lot of engineering behind making both true at once.
Thank you
@NVIDIA and the Nemotron team,
@OpenRouter and
@alexatallah for allowing providers like ourselves to work seamlessly, and our engineers, who know how much work sits behind a neat row in a pricing table.