The sea is the sea The old man is an old man The boy is a boy and the fish is a fish The sharks are all sharks no better and no worse

San Francisco
Today we're launching Base Labs, a research lab by Baseten. Our mandate is to make open-source as useful as possible, and our one rule is that we publish without exception, including what fails. Up until pretty recently I thought the way to get the world onto open models was to train them for one company at a time. @mudithj, @maxkirkby and I cofounded @parsedlabs on that bet, @baseten acquired us, and we spent the last year running their training team doing it for customers one by one. Every one of those engagements taught us something new about how models learn, forget, specialise and get cheaper, and almost none of it got spoken about. Sadly, in general that's the field's default in that the people who know the most about training have the least freedom to say it. To be clear I don't think the closed labs are the villains here. They get to new capabilities first, which buys the rest of us time to harden the world before that stuff is everywhere, and they're the ones paying to find out what's actually possible. But I'm fairly convinced the only real advantage they have is data and scale, and their incentives point squarely at the frontier. You can't do slow, public science on how these things learn when your job is the best model by end of quarter. Someone without that pressure has to, and there are very few of those someones around. Hence Base Labs. We have the broad remit of making open-source as useful as possible and our one rule is that we publish without exception, including what fails. The first problem is continual learning, which I have come to think is several problems wearing one name. We are also working on the open RL environments and data that open models need and currently can't get, because we have to aggregate data with the same ferocity everyone's been talking about aggregating compute. Plus a bunch of other stuff I'm genuinely excited about, eg a safety stack people can run on top of open deployments, and performance research so these things are cheap for everyone to serve. I still think open and closed coexist, and that's the good world. It's just that coexistence isn't free, someone has to actually do the work, and this is basically what keeps me up at night. We're hiring researchers, engineers and fellows. Come help distribute the mandate of heaven!
Today we're announcing Base Labs, a dedicated research organization focused on advancing open-source AI. We believe in a healthy, open frontier model ecosystem. To enable this, we are working on: - Blue-sky research on continual learning, the science of RL, and how models learn across their full lifecycle, with every experiment and recipe shared openly. - The BaseHub Data Foundry: the highest-quality open RL environments, training data, and real-world benchmarks, built for anyone to train and benchmark on. - Post-post training: taking open-source models and making them better, safer, and more aligned through continual post-training, built on our research, and deployed with our frontier safety stack so organizations can use open-source models with confidence. - Making models cheaper and more performant through our model performance research. This is a mission-driven research effort, not a commercial product. We believe the health of the open-source AI ecosystem matters and that the best way to advance it is to do science in the open. We’re hiring engineers, researchers, and research fellows to advance this mission. labs.baseten.co/
68
64
930
125,858
sachin and adam make the best videos. hopefully the our chats do their production quality justice
Exclusive: Inside @baseten The $13B startup at the centre of the biggest fight in AI - Open vs Closed. We spent a day at their SF HQ with co-founder & CEO Tuhin Srivastava, Charlie O'Neill, Mudith Jayasekara and the Base Labs team - to understand why they're building an open ecosystem. Backed by top VCs led by @saranormous @altcap @apoorv03 @wmccarthy11 @saammotamedi @nikiscevak @lachygroom @willreed + more. Plus participation from @nvidia. 0:00 The Open vs Closed AI War 2:18 A Billion AI Requests a Day 3:43 The Best Models Are Chinese 4:06 Tuhin: Don't Rent Intelligence 5:09 Sachin & Tuhin Talk Cricket 6:02 Tuhin: The Real Bottleneck 10:15 Why Launch Base Labs? 11:35 "You Have to Have an Opinion" 12:52 Centralised Control Is the Danger 13:07 Publish Everything, Even Fails 15:08 Hugging Face & the 6-Month Gap 16:10 Why We Still Need Closed
3
3
56
7,892
I really miss having smart haters 😭
2
19
2,961
Charlie O'Neill retweeted
OpenAI serving open-source models through Baseten is a big deal. They are embracing open source. Until now, OpenAI's pitch was "use our models." Now it is "commit your annual AI budget to us". We will get you the best model for each job, whether it's ours, someone else's, or open source. So if you're a large bank looking to commit $100M in AI spend, you centralize with OpenAI. One throat to choke. GPT handles the credit memos. Open-source handles the KYC volume at a fraction of the cost. A post-trained / specialized model handles fraud. This changes OpenAI's position. It stops competing only on whose model is best this season and starts competing to own the enterprise AI budget. A layer above being just an AI lab. Also proves AI is positive sum, not zero sum between closed and open source!
14
28
350
47,043
Charlie O'Neill retweeted
OpenAI <> Baseten 🚀 Customers can now use OpenAI commits towards open-source tokens via @baseten!
1
3
67
4,661
opsd and its variants don’t work, no matter what they are biased and bias kills llms, please please please stop working on it, so many other great directions to explore that aren’t opsd
Meta's team proposes DCE + SRCL as an alternative to on-policy self-distillation (OPSD). Achieved significant gains: 30.76% avg acc -> 65.97% on Qwen3-8B arxiv.org/abs/2609.30652
24
34
434
86,172
The sun has finally set on the British empire. This is genuinely so sad to me. The country that gave the world the steam engine is now treating the most revolutionary technology in history as a dirty, ugly tool. Straight out of atlas shrugged
BREAKING: the UK government publishes official rules on how to use AI: 1. Ask first if you really need to use AI 2. Check if a spreadsheet can do the job before using AI 3. Choose the worst model possible so it uses less energy 4. Keep prompts short to reduce the environmental impact 5. In general, use AI only when necessary, as a climate-saving measure With a mindset like this, we should accept the UK back into the European Union
35
61
1,029
112,476
in fairness to openai i think an xgboost model could hack medicare
Australia has been hacked. 'And today, I spoke with the CEO of OpenAI, Sam Altman, to express Australia's extreme concern about this incident. And I also expressed my disappointment that it took the company way too long to inform the government what had occurred, and the nature of the way that that notification occurred as well was unacceptable.'
2
5
178
13,680
Distillation is used to warmstart your model for RL. But there are many unanswered questions. What is the exchange rate of distillation and RL? Does distillation burn entropy and thus cap the model's ceiling? Are there behaviours you can only learn if you distill from a stronger model, that RL will never find? This is the first step towards starting to frame and answer some of those questions. Thank you to @dwarkesh_sp for many interesting discussions which motivated a lot of this research
Distillation is a common way to prepare a model for RL where learning occurs from a stronger model, and then improves through trial and error. But how much does that better starting point help after RL? We conduct preliminary research on this question across model sizes and reasoning tasks. 🧵
12
5
194
18,985
if this doesn’t wake you up to the fact that we have officially unleashed a new scaling axis I don’t know what will whenever log plots break along some stable metric but overall progress feels like it’s improving at same expected rate, this is a good sign you’ve got a new scaling axis. We’ve seen the same thing with Moore’s law over and over
> OpenAI be like: I swear the Navier Stokes run was not that bad > looks inside > 💀
9
16
370
72,217
Charlie O'Neill retweeted
Happy to start collaborating with @baselabs from @baseten and @GoodfireAI to push safety and interpretability for open models. As open-source models catch up in performance to frontier models, the community at large have a great opportunity to establish a common practice for effective safety, security and interpretability research. Some past discussions confused “open” with “unmonitored”. On the contrary, an open model provides much more tooling and visibility for safety research from a broader audience, which has led to a lot of the safety and security techniques we use today. Linux is a great example of this: open-source and secure deployment are not only compatible but heavily intertwined, as a properly secure system needs an extensive feedback loop of finding and fixing vulnerabilities. We're excited to share more soon.
Safety is not just for closed models. The closed frontier labs are a canary in the coal mine for what is coming at scale. They give us a glimpse into the future and a window to harden our systems and prepare for abundant intelligence, with all the risks that come along with it. The OpenAI agent swarm attack on Hugging Face is the kind of failure we need to prepare for as open-source models catch up. The providers serving those models (such as Baseten) have a big role in establishing what safety and monitoring standards look like. We’re proud to be taking the lead on this at @baselabs with our collaborators @huggingface and @GoodfireAI. We’re developing safety research in the open and building it directly into Baseten’s inference infrastructure, with the aim of making it available to all our customers. This includes training models to follow explicit policies, detecting failures at runtime, and connecting those signals to controls that can intervene. We invite others in the open-source ecosystem to join us in building the tools and standards we’ll all need.
14
13
151
20,575
Charlie O'Neill retweeted
Proud to partner with @baseten and @baselabs to build safety infrastructure for open-source models—which are essential to lots of safety research, including our own. Safety must be built into open models and provided by those who serve them, and we’re excited to help enable that!
We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source. This is why Baseten and Base Labs are building a stronger safety and security standard for open models, with the launch of our safety infrastructure. Base Labs will develop and publish methods for training and monitoring open models, and Baseten will integrate that work into its deployment infrastructure, live at runtime, and offer this work as a managed service. This will be a standard that is transparent and built into how our models are trained and deployed. We invite the open-source community to contribute, and are proud to partner with @huggingface and @GoodfireAI to bring this vision to fruition. Together, we are building an ecosystem of open models that are safe and accessible to all.
2
20
189
14,658
Safety is not just for closed models. The closed frontier labs are a canary in the coal mine for what is coming at scale. They give us a glimpse into the future and a window to harden our systems and prepare for abundant intelligence, with all the risks that come along with it. The OpenAI agent swarm attack on Hugging Face is the kind of failure we need to prepare for as open-source models catch up. The providers serving those models (such as Baseten) have a big role in establishing what safety and monitoring standards look like. We’re proud to be taking the lead on this at @baselabs with our collaborators @huggingface and @GoodfireAI. We’re developing safety research in the open and building it directly into Baseten’s inference infrastructure, with the aim of making it available to all our customers. This includes training models to follow explicit policies, detecting failures at runtime, and connecting those signals to controls that can intervene. We invite others in the open-source ecosystem to join us in building the tools and standards we’ll all need.
We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source. This is why Baseten and Base Labs are building a stronger safety and security standard for open models, with the launch of our safety infrastructure. Base Labs will develop and publish methods for training and monitoring open models, and Baseten will integrate that work into its deployment infrastructure, live at runtime, and offer this work as a managed service. This will be a standard that is transparent and built into how our models are trained and deployed. We invite the open-source community to contribute, and are proud to partner with @huggingface and @GoodfireAI to bring this vision to fruition. Together, we are building an ecosystem of open models that are safe and accessible to all.
4
14
131
29,983
and when you realise most rl envs are just LLM evals but slop volumed, a lot of things (eg misalignment) start to make a lot more sense
at some point we need to seriously have a discussion about the state of LLM evals. reading the traces and seeing truly horrible stuff
9
23
405
25,080
This is the last time you’ll be able to spot this. It’s like AI generated images start of this year. You’ll stop spotting it, because people will get so good at it. That is, because every single product, experience, game, thing will become so well engineered at capturing real world inputs to seed RL envs / midtraining data. There will be whole design philosophies on doing this without the user realising. Dark user patterns will abound. The world will become a factory for RL envs
never seen a more obvious raytheon psyop to farm training data
7
18
283
37,022
very kind from AT but most importantly I agree that someone should figure out the contours of this conflict: "on one hand: <> “anything that can be learned through RL can be distilled very easily”. this is the justification as to why open source models (chinese) can catch up to closed source models (american). fair... 5 minutes later <> “Opus 5 could not distill / generalize well from Fable [despite anthropic having the live deployment, access to logits, and obviously a good prompt distribution]”." we are working on a version of this at @baselabs
extremely high snr from this podcast...the only one i've been able to watch in full in one sitting. <> the only way we don't get rsi is if we fall into some regulatory capture (which seems to be trending at present) <> we're nowhere near the ceiling of how well you can do research <> all thinking can do is update your posterior based on the knowledge you’ve gained since you formed your prior. you can’t gain any new knowledge from just thinking <> you can spend an equivalent amount [7 figures] of compute in AI agents to get a century’s worth of thinking, a century's worth of theory, before every training run <> taste is just behavior that works in the long run, and can be baked in a longer context window <> creativity is just solving hard search problems, and can also be baked in a longer context window you should follow everyone here, especially @oneill_c, i think it takes a special talent to be able to not only develop deep technical competency, but to also be able to use that to consistently make accurate predictions about the future (i think this was literally François Chollet's definition of intelligence in last year's YC event). my only nit here: @dwarkesh_sp should have pushed on two quite conflicting statements from @oneill_c and @BerenMillidge. on one hand: <> “anything that can be learned through RL can be distilled very easily”. this is the justification as to why open source models (chinese) can catch up to closed source models (american). fair... 5 minutes later <> “Opus 5 could not distill / generalize well from Fable [despite anthropic having the live deployment, access to logits, and obviously a good prompt distribution]”. can only have one or the other, imo. also lol at the timeline: - ai will dominate top human experts in 3-4 years - it will take 5-10 years to automate ai research k i n o
1
31
4,571
I can guarantee that google is behind OpenAI and Anthropic in terms of rsi
Altman, Amodei and Musk are likely realizing that Google did not focus their external efforts on improving consumer ai and instead went all in on unlocking ASI.
13
128
28,310
I think a downstream consequence of increased commitment to safety and alignment is that RL envs companies get screwed, or at least held to a much higher standard and thus their unit economics changes. The evidence is mixed on this early but it seems likely that impossible to poor quality envs lead to misalignment and poor behaviour in the models (as they try to do anything to get reward). So why would the labs risk this quality control with anything but big in-house efforts moving forward
8
5
180
20,778
Charlie O'Neill retweeted
A flavour of our research interests on the Dwarkesh Podcast this week, discussed by our very own @oneill_c. Just the start of a longer conversation about long-horizon RL and frontier open-source training here at Base Labs.
New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
1
3
25
3,848
For better or for worse, chamath + All In pod is somewhere that so many people, particularly non technical, get exposure to information and opinions on AI. Right now is a perfect opportunity to get more informed and thoughtful voices on there
Dario is making the case for the opposite. This actually makes our life harder and makes it easier for others to catch up with us, but we still think it is the right thing to do. Happy to come on the pod next week and talk about it!
5
3
200
35,795
Charlie O'Neill retweeted
really cool, didn't know @oneill_c was Australian until i heard the accent I hadn't really considered the prompt distribution and the usefulness of realistic data in that regard and how the distillation goes through these channels, its a good explanation of why Chinese models are so good compared to like sonnet or opus though.
New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
1
6
1,519