The best AI is built, not bought. Our platform, Applied Compute Agent Cloud, is now in private beta. Book a demo below.

San Francisco
Today, we're introducing AC2, the Applied Compute Agent Cloud, to enable every team to train, serve, and improve their own frontier models.
30
39
456
205,199
Our reward hacking monitor is how our Applied Researchers, like @evandavidellis, catch models cheating during RL. This is calibrated on an internal benchmark of real reward hacks we've observed in our runs. Today, we're rolling out this monitor into AC2.
We're adding reward-hacking monitors into AC2. Every RL run gets checked for cheating, and our research agent, Ari, automatically investigates the flags. After a model discovers an exploit, it can become its default strategy within a few training steps. We used to hunt for reward hacking manually, reading traces and running one-off agent queries. This was error prone and didn't scale. We decided to automate detection instead.
1
2
41
2,843
Applied Compute retweeted
Next up on First Pass: A conversation with @ypatil125 from @appliedcompute. Customized models are coming. But how can enterprises get ready for this future? We dug in on questions like that with Yash
7
5
67
10,977
Applied Compute retweeted
I asked people: "where should enterprises go for post-training as a service?" 30 votes: Applied Compute (@ypatil125, @rhythmrg, @lindensli) 13 votes: Trajectory (@rronak_, @michaelelabd, @QuantumArjun) 10 votes: Prime Intellect (@vincentweisser, @johannes_hage) 7 votes: Thinking Machines (@miramurati, @johnschulman2) 5 votes: Fireworks (@lqiao, @dzhulgakov, @pawelg, @jamesr66a, Chenyu Zhao, @divchenko) 4 votes: Baseten (@tuhinone, @amiruci, @saltyph, @defpan, @mudithj, @maxkirkby, @oneill_c) 4 votes: Together AI (@vipulved, @ce_zhang, Chris Ré, @tri_dao, @percyliang) 2 votes: Arcanum (Rehan Rupawalla, @danielrupawalla, Krish Thawani, @samir_walji) 2 votes: Belvedir (@thezacharyyu) 2 votes: Trainloop (@jackson_stokes, @mlpierce22) 1 vote each: - AfterQuery (@CarlosGeorgescu, @spencermateega) - Bespoke Labs (@madiator, @AlexGDimakis) - Engram (@dan_biderman, @EyubogluSabri, @realJessyLin) - Mercor (@BrendanFoody, @adarsh_exe, @suryamidha) - Radixark (@ying11231, @BanghuaZ) - River AI (@ibab, @dsoboliev, Ievgen Soboliev, Yaroslav Nazarov) - Unsloth (@danielhanchen, Michael Han) People couldn't pick a company they work at or founded. Disclosure: I'm a small investor in AfterQuery, Applied Compute, Arcanum and Trajectory. I didn't vote.
14
28
417
47,291
The adoption of open weights is accelerating, but how do we securely deploy them? We've assembled a framework for assessing and mitigating risks in open weight model deployments based on our experiences post-training with frontier enterprises.
4
1
40
3,919
Gap analysis is how our Applied Researchers, like @_brylee10, surface agent failures across billions of tokens of RL traces for our customers. Using our platform, AC2, and low-cost classifiers like Jev, we can catch 85% of failure modes at a fraction of the cost of LLM judges and turn them into training data.
I implemented a system in @appliedcompute’s platform for automated failure mode clustering with Jev to surface errors at an even larger scale than before. RL training produces billions of tokens in traces. I always manually read many traces to understand model behavior, but finding agent failures (like reward hacking / hallucinations) at scale is easy to miss without automation. Here’s how it works:
3
2
38
4,759
Applied Compute retweeted
we partnered with @appliedcompute to post-train a small model for large-scale code search over precomputed indexes at 300 repos, this is ~3x faster than using filesystem + grep, and reduces the marginal cost of a search by up to 100x vs frontier models turbopuffer.com/blog/large-s…
10
39
437
40,890
A 35B open-weight model trained to search a precomputed index answers repo search questions at 100x lower cost than a frontier model. We partnered with @turbopuffer to train Qwen3.6-35B-A3B to find code across ~9,000 repositories. It tops the needle-in-a-haystack task outright at 2-10x lower latency.
6
30
354
48,405
Visit codesearch.appliedcompute.co… to watch our post-trained agent search 2,789 repositories with cited results, including open-ended asks like finding fun ASCII art.
1
23
1,238
Out of the box the base model is a weak searcher. Over training, correctness x citation support improves 57% on the narrow task and precision improves 211% on the open-ended one. The agent also learns where to look. ripgrep falls from 15.3 calls per rollout to 0.3 as turbopuffer search takes over.
92
Applied Compute retweeted
Agents are the new discovery layer. We've seen that across 2M+ conversations with Ask DoorDash, where people want agents to step in and carry their intent through to a fulfillable order.
“50% of DoorDash’s agentic restaurant orders are going to places users have never ordered from before.” @andyfang tells our CEO @ypatil125 what happens when agents become the discovery layer. If models increasingly decide what gets surfaced and bought, companies have a strong reason to train and own that intelligence.
1
4
44
8,362
“50% of DoorDash’s agentic restaurant orders are going to places users have never ordered from before.” @andyfang tells our CEO @ypatil125 what happens when agents become the discovery layer. If models increasingly decide what gets surfaced and bought, companies have a strong reason to train and own that intelligence.
3
25
11,812
“Until you actually see things operationally, it’s going to be hard to build DoorDash from scratch.” Our CEO @ypatil125 sat down with @andyfang on why cheaper software doesn’t erase years of operating advantage. @DoorDash’s moat is its proprietary data, edge cases, and hard-won knowledge, and increasingly, the models trained on top of it.
5
42
12,342
Applied Compute retweeted
AFAIK AC2 Is the only platform with support for full weight fine tuning of Kimi K3. For short, simple tasks LoRA fine tuning may be similar. But for multi turn, agentic workloads, full weight training is strictly superior. dm if you’re interested in trying AC2
Kimi K3 full fine-tuning is live on AC2. Our memory optimizations reduced GPUs required per training replica by ~40%. At nearly 3T parameters, Kimi forced us to rethink how we manage memory, communication, rollouts, and checkpoints. The result is a much more efficient path to training frontier-scale open models.
2
20
3,004
Kimi K3 full fine-tuning is live on AC2. Our memory optimizations reduced GPUs required per training replica by ~40%. At nearly 3T parameters, Kimi forced us to rethink how we manage memory, communication, rollouts, and checkpoints. The result is a much more efficient path to training frontier-scale open models.
3
51
647
235,827
On the inference side, MXFP4 rollouts ran on 2 B300 nodes instead of 4, while staying within the same KL range we observed with bf16 inference. Less memory and fewer inference nodes directly lower the cost of frontier-scale RL.
1
2
39
4,669
Supporting Kimi K3 required changes across training memory, rollout precision, weight transfer, checkpointing, and communication. Those improvements also carry over to other models on AC2. Read the full report. appliedcompute.com/research/…
3
46
3,571