lead product for inference @togethercompute. invest in ai/ml, infra, dev tools. ex @temporalio, @google

San Francisco, CA
i'm hiring! job-boards.greenhouse.io/tog… i would love to talk to folks who geek out on APIs, agent experiences, harnesses, and open-weight models
16
14
193
24,157
Nikitha Suryadevara retweeted
2 yrs ago i had voice agents all around to all the coffee shops in sf and ask if they have ac it’s live again, updating it with corgi data now peytoncasper.com/ac-map/
Saturday will be one of the hottest days of the year for the entire Bay Area. Downtown San Francisco is forecast to reach 94°F in the afternoon, while Ocean Beach, just over 6 miles to the west, is expected to be in the low 60s. The Outer Sunset and Outer Richmond could briefly climb into the 80s around noon before the sea breeze kicks in during the afternoon. Welcome to Indian Summer.
2
1
17
1,568
Nikitha Suryadevara retweeted
Today we're excited to announce @HalluminateAI's $30M Series A led by @oakhcft with participation from @YCombinator, @orangecollectv, @FTPartners, @heavybit, and more.
51
23
223
47,109
Nikitha Suryadevara retweeted
shhh today we’re launching function secrets our customer’s agents are performing millions of logins every day and they should have a secure way to handle them
Browserbase Functions now support Secrets. Agents need credentials to do real work on your behalf. With Secrets, you can securely store credentials on the Browserbase platform and reference them at runtime in Functions. Give agents access without leaking secrets.
3
31
1,888
i didn’t realize how much working on k8s/borg would prepare me for running these beasts
Our serverless APIs for open weights models have become exquisite machinery (that we plan to write more about soon!) One effect is that @togethercompute now one of the largest originators of open tokens in the US, with a wild week over week growth curve that reflects the pace of OSS adoption. Will serve 1.8T tokens today on @OpenRouter alone. You can get started with $5 — api.together.ai openrouter.ai/provider/toget…
10
668
Nikitha Suryadevara retweeted
are you seeing a pattern here? > GLM 5.3 > GLM 5.3-Flash > DeepSeek 4.1 Flash > Kimi K3 top agentic models token factory going brrrr
6
3
18
2,225
pro tip: flour + water does half portions of most things on their menu if you go solo. perfect portion sizes!
7
380
Nikitha Suryadevara retweeted
Our dedicated model inference @togethercompute now supports canary rollouts with gated steps for safe, cost efficient and zero downtime upgrades on live traffic. A post on how to use this feature.
New on Dedicated Model Inference: canary rollouts. Upgrade the model behind a live endpoint without downtime. Traffic moves from your current deployment to the new checkpoint in gated steps (default 5% → 25% → 50% → 100%). Health checks run before any traffic shifts. After every step, metric gates compare the new model's p95 latency and error rate against the old one. If a gate trips, the rollout pauses at the canary share and waits for you: resume, promote to 100%, or roll back. Three strategies: canary, blue-green, and rolling. Available now via the tg CLI, REST API, and Python SDK. Learn how to start a rollout: together.ai/blog/canary-roll…
3
3
17
2,593
i did not realize growing up to be a real adult would mean getting an inordinate amount of joy from replacing my regular toilet with a toto
4
293
Nikitha Suryadevara retweeted
omg @togethercompute I love you. this is time to first token for every provider for Deepseek v4.1 flash. you guys CRUSH on pre-warmed queries. there's no other game in town
2
7
15
7,047
Nikitha Suryadevara retweeted
We did it again. Temporal just raised a $550M Series E round at a $12.55B valuation. Learn more about how we got here: temporal.io/blog/temporal-ra…
53
49
744
208,803
Nikitha Suryadevara retweeted
Nice stats from @OpenRouter that illustrate how @togethercompute delivers solid and scaled performance for agentic workloads. We are serving GLM 5.3 and GLM 5.3 Flash @ top decile of TPS, latency, cache rate, and doing it at large volumes, 23% and 30% of all OpenRouter traffic, and OpenRouter is a fraction of our overall API traffic. Running scaled inference for agentic applications is a lot more than maxxing a single metric ... you need to optimize on all dimensions and doing so at scale with reliability. openrouter.ai/z-ai/glm-5.3 openrouter.ai/z-ai/glm-5.3-f…
7
11
61
7,860
Nikitha Suryadevara retweeted
i put cad in a terminal, won @OpenAI gpt6 hackathon and $50k cad sandbox puts fusion 360 on a vm, wrapped its sdk with a server, live streamed model state to a browser running in a terminal and let astra go nuts $5k in tokens to the craziest idea in my dms and a shot for every mech e that’s been personally attacked by this idea
Replying to @AlexReibman
3/ CAD Sandboxes 🥇$50,000 credits + DevDay tickets SDK and visual design inspector for Astra to interface with Fusion 360 via Codex Browserbase, but for CAD @peytoncasper
11
7
55
3,547
Nikitha Suryadevara retweeted
our conference on the state of the agentic web + computer use could not have been better timed! join us in SF on thursday for a glimpse into the near future: browserbase.com/navigate
It's Navigate week! Browserbase is hosting Navigate this Thursday (Sept 10) in SF. People are flying in from across the US to discuss the future of agents and computer use. Listen to speakers from @stripe, @vercel, @AnthropicAI, and more. Seats are running low. If you want to join the conversation on the future of agents, registration officially closes tomorrow.
3
5
51
7,526
Nikitha Suryadevara retweeted
ramp quietly shipped one of the best agent identity implementations on the web @shreypandya built a procurement agent that can log into basically any vendor, pull the receipt, and upload it to ramp using their new agent identity agent identity and commerce are coming in hot
Agents need their own identity to do real work on the web. We partnered with @tryramp to automate our event logistics workflow. Our agent logs into sites with a Browserbase Context, downloads receipts, and submits them via the Ramp CLI.
3
1
64
10,081
we shipped this baby 6 months ago and how it has grown 🌱 dedicated container inference is the best place to run diffusion models. super excited to work with @higgsfield_ai!
Huge welcome today: @higgsfield_ai, the leading AI video and image creation platform behind Cinema Studio, is now a Together AI customer, fresh off their $400M Series B. Cinematic intelligence, one unified workflow, studio-quality video at any scale. Their video models run on Dedicated Container Inference, built for long-running, multi-GPU jobs: autoscaling, queues, traffic isolation, retries, monitoring. Proud to power the inference behind it.
1
23
2,077
Nikitha Suryadevara retweeted
10,000 @nvidia B300 GPUs. India's largest AI Factory. Together AI and @larsentoubro are building the country's biggest GPU cluster, backing open-source inference, fine-tuning, and training at scale for India's AI-native ecosystem. reuters.com/world/india/indi…
7
21
201
45,244
Nikitha Suryadevara retweeted
hello world, meet sigil, a new way to write for work. it has been a joy to work on something i care about deeply with @jackiehluo. we first met as 19 year olds in manhattan obsessed with writing and startups. not much has changed, except we're in brooklyn now.
introducing sigil. it's a new way to write for work that gives you the power of machine intelligence while putting your words and ideas first. here's why we think that matters. sigil.so
33
10
179
31,224
Nikitha Suryadevara retweeted
Frontier performance you can actually own. Proud to help power DeepSeek-V4-Flash on Ollama's cloud, with the fastest hosted performance available. Open weights, zero data retention, U.S. & EU hosting.
DeepSeek-V4-Flash-0731 is now fully rolled out as the new default for deepseek-v4-flash on Ollama's cloud. This model combines speed, efficiency, and frontier-level performance. Fast: 120+ output tps on Ollama's cloud Private: zero data retention hosting in US & Europe Efficient: generous usage on Ollama's Pro and Max plans for multiple long-running, uninterrupted sessions with your favorite coding harnesses.
10
12
61
14,517