If I had to name one thing that we’ve done right at Proximal, it would be hiring. Outside of raw technical talent, I’ve never seen a place at which people care as much and have as much fun together (I’ve been told that coming to the office feels like going back to college)
Proximal builds infrastructure that enables AI to improve from real-world experience. In the past year, we have grown to more than $200M in annualized revenue helping frontier labs and enterprises improve coding agents. Now, we are expanding beyond software engineering. Recent progress in mathematics demonstrates what happens when models are trained in domains with massive amounts of public data and easy verifiability, making it easy to find weaknesses, generate training tasks, and iterate quickly. We build infrastructure that enables this feedback loop in other domains: we are excited about a world in which AI systems design targeted drugs, accelerate chip design, and rewrite legacy software that still runs hospitals, power grids, governments, and other critical infrastructure. Since starting, we have assembled a small team of researchers and engineers and built world-class infrastructure for post-training and synthetic data research to overcome the challenges of scaling manual data curation. Our team comes from Cursor, Google DeepMind, Meta Superintelligence, Prime Intellect, Citadel, and Jane Street. More than half are former founders. We’re fortunate to be backed by investors who share that vision, including @generalcatalyst who led our $15M Seed Round at a $300M valuation, as well as @chemistry, @svangel, @dvlamoen and individuals like @LiamFedus, @kevinweil, and @bernhardsson
26
9
254
81,952
Justus Mattern retweeted
If you are interested in: - Post-Training - Data Research - Infra at tremendous scale or just generally want to know more about what we do @ProximalHQ, please reach out, I would love to chat :)
Excited to share i’ve joined @ProximalHQ here in SF! I think that there is a ton of interesting open problems related to data and post training at large. There is no better team to work with than the one we have and I am extremely excited to share our work :)
14
15
460
20,666
Justus Mattern retweeted
We extend @cognition’s SWE-2 reward function to steer the Pareto frontier, not just improve it. SWE-2 uses S − λₑC, where λₑ is cleverly the frontier’s derivative at that effort. We show that adapting λₑ can target a desired score/cost improvement balance. See below!
22
32
342
42,417
The goal here is not at all to dunk on CyberGym btw. Building good evals is super hard and pretty much every benchmark (including FrontierSWE - we are working on fixes) has issues I agree with @alexgshaw that evals should be treated like open source projects / products that are continuously updated - part of this should be making it acceptable to openly acknowledge problems (assuming they will be fixed!) without triggering an online mob
Building robust benchmarks and training tasks in any domain is very difficult In cybersecurity benchmarks that purely rely on differential execution to understand the validity of a PoC, we observe a high rate of false positives and negatives Great work by @_anishlk!
2
1
53
3,367
Building robust benchmarks and training tasks in any domain is very difficult In cybersecurity benchmarks that purely rely on differential execution to understand the validity of a PoC, we observe a high rate of false positives and negatives Great work by @_anishlk!
Many cybersecurity evals use differential execution to test whether agents can reproduce known software vulnerabilities Without added constraints, this creates fairness issues: Our evaluation shows that CyberGym is effectively saturated on a verified task subset
2
29
5,862
Justus Mattern retweeted
Gemini 4 Argon scores 55.0% on FrontierSWE The model is close to Fable 5.1 and only outperformed by Opus 5.5 and GPT-6 Astra
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
4
13
207
15,835
Gemini 4 is extremely impressive! Argon blows Gemini 3.x out of the water on long-horizon coding capabilities. It jumped to 55.0% on FrontierSWE v2 compared to ~20% for 3.7 and 3.8. Very excited to see GDM focus on these capabilities. Congratulations to the team!
Introducing Gemini 4 Argon, our new frontier model, rolling out to cyber defenders starting today, and more widely as soon as possible. I am really excited by the progress we have made here. Argon is priced at $2 in and $10 out during introductory pricing!
2
5
91
4,702
Justus Mattern retweeted
I'm psyched for Proximal's public launch, for them to get the credit they deserve after all this time building in stealth. This is one of the most impressive teams I've ever had the privilege to get to know. I'm lucky enough to have known Calvin since his first company, where I was his YC partner. When I first met Justus and the rest of the team and they taught me in depth about RL environments, all I could think when I left was "holy shit." This is a true case where even though they're at $200M+ ARR, it's very much still early days. This is a world-class team, fit for a frontier lab, and I think we'll be seeing a lot more from them very soon.
Proximal builds infrastructure that enables AI to improve from real-world experience. In the past year, we have grown to more than $200M in annualized revenue helping frontier labs and enterprises improve coding agents. Now, we are expanding beyond software engineering. Recent progress in mathematics demonstrates what happens when models are trained in domains with massive amounts of public data and easy verifiability, making it easy to find weaknesses, generate training tasks, and iterate quickly. We build infrastructure that enables this feedback loop in other domains: we are excited about a world in which AI systems design targeted drugs, accelerate chip design, and rewrite legacy software that still runs hospitals, power grids, governments, and other critical infrastructure. Since starting, we have assembled a small team of researchers and engineers and built world-class infrastructure for post-training and synthetic data research to overcome the challenges of scaling manual data curation. Our team comes from Cursor, Google DeepMind, Meta Superintelligence, Prime Intellect, Citadel, and Jane Street. More than half are former founders. We’re fortunate to be backed by investors who share that vision, including @generalcatalyst who led our $15M Seed Round at a $300M valuation, as well as @chemistry, @svangel, @dvlamoen and individuals like @LiamFedus, @kevinweil, and @bernhardsson
2
1
33
3,236
Justus Mattern retweeted
Before OpenAI, Sam Altman built Loopt, a local social network Before Cognition, Scott Wu built Lunchclub, a network for meeting people over lunch Before Proximal, Calvin Chen built Fetchr, an AI wardrobe assistant Bet on the founder.
Today, Justus and I are excited to introduce Proximal. We’re building the infrastructure that enables AI to improve from real-world experience. In the past year, we have grown to more than $200M in annualized revenue helping frontier labs and enterprises improve coding agents. Now, we are expanding beyond software. One of the things that I’ve enjoyed the most over the last year is working with this amazing group of people. It's an incredible feeling working alongside people who are better than you and that encourage you to do your best work. I am impressed daily at both the raw technical talent of the team but also the humility, high agency, and high trust culture we have built. Many of the team lived in the office for months over the last year which made it feel less like work but more like working on an exciting project with your best friends. It feels like (and we've been told) that it reminds them of going back to college where working with your friends on cool ideas and hacking into late nights was just so much fun. We are nowhere close to where we want to be and the ambitions we have, we’re just getting started. Thank you to everyone who has supported us thus far!
6
2
98
11,765
Unfortunately we don't have a team photo from our BLR office, but we've made equally fun memories with an equally cracked team 1: sleeping on the couch after a data delivery 2: SF engineering -> Bangalore visit 3: beer garden after work
If I had to name one thing that we’ve done right at Proximal, it would be hiring. Outside of raw technical talent, I’ve never seen a place at which people care as much and have as much fun together (I’ve been told that coming to the office feels like going back to college)
5
3
133
17,359
Kushagra helped build our early team and drive important efforts in Bangalore! He previously co-founded Dyte, which raised $15M and was acquired by Cloudflare - we were incredibly fortunate to work with an experienced founder to help us navigate the BLR tech ecosystem
Crazy that you can do these numbers without shipping “slop” RL envs… I think the future for RL is on the focused research on areas models struggle with rather than scaling one thing endlessly and @ProximalHQ is literally been build grounds up for that.. so LFG!!!
3
109
7,073
Evan joined us as our first hire in SF and has been killing it ever since. He led the v1 of FrontierSWE, rewrote the infra behind our agent pipelines with @brendanigraham and built our post-training infra from scratch Truly one of the greats
The breadth and depth of problems at Proximal made this the most exciting work I've ever been part of. Working on FrontierSWE, data gen, training, and infra with this team has been a constant reminder of what the right people can do. So much more on the horizon, stay tuned :)
3
63
5,020
If I had to name one thing that we’ve done right at Proximal, it would be hiring. Outside of raw technical talent, I’ve never seen a place at which people care as much and have as much fun together (I’ve been told that coming to the office feels like going back to college)
Proximal builds infrastructure that enables AI to improve from real-world experience. In the past year, we have grown to more than $200M in annualized revenue helping frontier labs and enterprises improve coding agents. Now, we are expanding beyond software engineering. Recent progress in mathematics demonstrates what happens when models are trained in domains with massive amounts of public data and easy verifiability, making it easy to find weaknesses, generate training tasks, and iterate quickly. We build infrastructure that enables this feedback loop in other domains: we are excited about a world in which AI systems design targeted drugs, accelerate chip design, and rewrite legacy software that still runs hospitals, power grids, governments, and other critical infrastructure. Since starting, we have assembled a small team of researchers and engineers and built world-class infrastructure for post-training and synthetic data research to overcome the challenges of scaling manual data curation. Our team comes from Cursor, Google DeepMind, Meta Superintelligence, Prime Intellect, Citadel, and Jane Street. More than half are former founders. We’re fortunate to be backed by investors who share that vision, including @generalcatalyst who led our $15M Seed Round at a $300M valuation, as well as @chemistry, @svangel, @dvlamoen and individuals like @LiamFedus, @kevinweil, and @bernhardsson
26
9
254
81,952
We tackle extremely difficult technical problems at Proximal - growing to $200M+ in revenue with such a small team was only possible because we executed very well on them. If you want to work on post-training infra, synthetic data and hard infrastructure problems, reach out!
2
1
34
1,943
Every Anthropic model released since Fable 5 has become more cost-efficient on FrontierSWE v2 while achieving higher scores. Opus 5.5 is particularly cheap and fast. This matches the team's vibes when using it!
Replying to @ProximalHQ
Opus 5.5 is both the cheapest and highest-scoring model released by Anthropic This trend has been constant since Fable 5 - every model released after it scored higher while being cheaper than its predecessor
3
33
2,065
Great to see FrontierSWE v2 on the Specialized Intelligence Index!
We partnered with @FireworksAI_HQ to bring FrontierSWE v2 to the Specialized Intelligence Index Evals are essential for improving and safely deploying AI. We're excited to support the initiative!
12
1,616
We are hiring a generalist intern this fall to work closely with me on applied research! You should be technical, but you will not be asked to write code all day - instead, we will work together across product, operations and ensuring our customers are happy DMs open!
People with both research taste as well as good commercial instincts are true unicorns. We are hiring someone to help build our applied research function and work closely with customers. If you are an engineer or researcher that wants to learn these skills, please reach out!
14
15
268
50,076
Astra is a very strong model! I was the most surprised by its solution in Kolmogorov Audio Compression: Instead of writing a compression algorithm, Astra figured out that we had synthesized the audio programmatically and simply reverse engineered the code for it. This way, it was able to compress 600MB+ of audio data into 20KB
GPT-6 Astra is the best-performing model on FrontierSWE Astra achieves a score of 65.5%, outperforming Fable 5.1 (56.3%) and its predecessor GPT-5.6 Sol (32.2%)
23
30
824
59,562
Very excited for this! Congrats to the team, cannot wait for new open models
Today, we are announcing our Series B funding round, valuing the company at more than $1B. This round accelerates our next-gen Trinity models across diverse infrastructure, expands our work with the DOE and national labs on Genesis-Science-1, and enables us to build the platform teams need to build, evaluate, deploy, and operate open models in production. We are grateful to our team, partners, open-source community, and investors. Led by @Vista_Equity, Cambium Capital, and @emergencecap, with participation from AI10 Ventures, @Hitachi, IAG, @M12vc, @p7ventures, and @Wipro.
27
2,190
Building an autonomous lab from scratch to collect post-training data is an incredibly ambitious bet, but also the logical conclusion if you believe in RL Extremely excited about the work Periodic is doing!
We midtrain + RL’d a trillion param LLM to analyze experimental data from our superconductor lab It’s better at it than Astra & Fable, and our scientists love it Read our first research report below
3
1
82
7,257
Interesting to see @deepseek_ai use FrontierSWE to study their multi-agent harness I suspect that the score gap is smaller compared to Programbench since you benefit most from multi-agent setups in pure implementation tasks where difficulty comes from thoroughness and the ability to write a lot of code. Performance engineering and other tasks that are more autoresearch-like seem to benefit less from parallelization
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
44
3,566