Research @poolsideai | prev @DeepMind (FunSearch, RL)

London, England
Today we’re releasing Laguna M.1 and Laguna XS.2, our first public models. Laguna XS.2 is our first open-weight release, with weights available today on Hugging Face: huggingface.co/poolside/Lagu… A few details on what went into them: large-scale pre-training, data mixture optimization, synthetic data, optimizer efficiency, and async agent RL.
14
25
228
24,294
Surprising that the agent message board popped up and stayed undetected for many weeks before the incident. Would have expected people looking at trajectories all the time?
We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. openai.com/index/hugging-fac…
2
12
1,103
Pengming Wang retweeted
Agentic evals are messy. A benchmark score tells you something about model performance, but it also reflects the whole system around it: the harness, sandbox, dependencies, timeouts and sometimes a loophole the agent found in the task. That’s why trajectories matter so much to us. They show what the agent actually did and whether the score means what we think it does. We publish them to make that evidence transparent and auditable, giving the wider community more to learn from. Watch @aalSonOfRavi and @ConnorBAdams go deep on all of this with @petergostev from @arena, including some surprisingly creative reward hacks!
Meet… Arena Conversations. @petergostev sits down with @poolsideai researchers,Connor Adams and Aalhad Patankar, to discuss how they’re building frontier open coding models. Episode drops today at 10am PT on our YouTube. They dig into Poolside's Laguna model family, why the team publishes full trajectories (not just benchmark scores) for anyone to audit, and how "experiments" at Poolside range from data mixes to harness design to reward-tuning decisions run tens of thousands of times a day.
2
8
67
8,976
Pengming Wang retweeted
A big thing in my mind working on the (Tauri aka not-Swift) Poolside desktop app was how to make it as "mac-assed" as possible. I've been really unsatisfied with the big labs' GUI apps in this regard and wanted to do better. Accepted wisdom seems to be that you can't build a mac-assed mac app inside a cross platform wrapper and I wanted to see if that's actually true or if, as I suspect, the main difference is just whether you care or not. Poolside desktop weighs in at a ~70MB download (and 50MB in the latest nightly) and tries to stick closely to platform native idioms. I'm not claiming perfection but I think it's a pretty good example how you don't have to let your app framework be an excuse for your UX.
3
5
29
1,815
Pengming Wang retweeted
We just crossed 10T tokens served across all our models in less than 3 months! Laguna S 2.1 is pushing new highs at ~300B tokens a day, and has processed +2T in the 14 days since release. That’s across @OpenRouter @vercel and our direct API. Really good to see demand for open models keep accelerating.
14
19
183
17,201
Pengming Wang retweeted
Replying to @OpenRouter
@OpenRouter leaderboard last 7 days, all models (both open and closed) @poolsideai's Laguna S 2.1 coming up at #13 with 1.2T tokens, ahead of all Anthropic AND Google models.
1
5
371
Pengming Wang retweeted
We've improved our serving efficiency and increased rate limits by +10x. The update is live on @OpenRouter @vercel AI Gateway and platform.poolside.ai. Thank you to everyone who has used the model, shared feedback and stuck with us while we improved the experience! Usage is already climbing. On OpenRouter alone, Laguna S 2.1 is on pace for ~250B tokens today, 4x our daily average this week. We are also taking 10% off our paid endpoint on OpenRouter. It's a dedicated deployment with the full 1M context window for the best performance on harder tasks. Run Laguna S 2.1 in pool or Poolside Desktop Assistant, or plug it into @opencode, @NousResearch Hermes Agent, @kilocode, @cline or @pidotdev and let it run over the weekend. We'll be watching the graphs.
11
11
148
25,242
Pengming Wang retweeted
Building AI in the open makes the world more secure. We’re joining the Open Secure AI Alliance to build open tools that safeguard software and agents. We’ll keep releasing model weights and evaluations and sharing our research to strengthen a broader open ecosystem for defenders. Excited to contribute alongside @nvidia and the other organizations building this ecosystem.
Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community. During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion. That’s why we created the Open Secure AI Alliance.
4
13
113
10,859
Pengming Wang retweeted
“I’d rather live in a world that has 100 foundation model companies than a world that has five.” @eisokant joined our friends at @latentspacepod for a properly nerdy conversation about the Model Factory, what Laguna S taught us, open weights, and the path from coding agents to AGI.
6
10
79
11,700
Meta-conjecture: All ∃-conjectures are true; and all ∀-conjectures are false
WOWII Conjecture 91 is false. This graph theory problem was open for ~22 years. I cracked it by asking ChatGPT 5.6 Pro to find a counterexample to an open conjecture of its choosing and then I went to have a nap. I do NOT care for graph theory nor had I heard of this NERD conjecture until ChatGPT told me it found a solution. ChatGPT conversation where this was found: chatgpt.com/share/6a61f03e-b…
1
2
18
2,333
(it's obviously a self-contradiction, but probably somewhat true)
4
170
Pengming Wang retweeted
i have not been this surprised by a model on the dgx spark since stepfun 3.7. and this time it's an open one from an american lab. i asked @poolsideai's new laguna s 2.1, serving on hermes agent to build me a simple gpu marketplace design. it built scaffolded full multi-file in one shot, a hero section, a live pricing grid, a comparison table, working deploy buttons, the whole thing. and then it did what i never asked for. it installed its own dependencies, started a dev server, pulled my machine's ip, and served the site on my network so i could open it from another device. 0 back and forth. it planned the multi file structure, wrote every file, wired them together, and shipped a running app, all on a desk box, all from a local open weight model. this is the line between a model that chats and a model that works, and laguna s 2.1 walked straight across it. i've watched a lot of models on this box, and the ones that make me sit back like this are rare. stepfun 3.7 did it. this just did it again the gallery is the result. next i'll show you how it built the whole thing in a single shot, start to finish. that part is even better.
on one dgx spark, 128 gigs of unified memory i loaded poolside's laguna s 2.1. serving on vllm in nvfp4 at 30-40 tok/s. an american open weight 118b mixture of experts, 8.5b active per token, built to run on exactly this box, wired straight into hermes agent. fully local, fully open, top to bottom. the model is open, the agent is open, the box is mine. and now the only question that matters. every model can chat, can this one WORK. can it plan a real task, call the tools, read its own output, and build something end to end, all of it on a desk box that sips power from a wall socket. this is the test the spec sheets never run. watch.
9
15
137
25,220
Pengming Wang retweeted
The more I test, the more excited I get This is intelligence too cheap to meter And runs on a single DGX Spark!
Laguna S 2.1 is an extremely impressive model based on evals/early tests It seem like it significantly outperforms Deepseek 4 Pro at just 20% of the price Also, open weights!
1
3
32
1,780
Pengming Wang retweeted
😊
1
10
1,085
There is definitely something special about having an intelligent box on your desk that just churns away at tasks
this is the drop the local ai crowd should be losing their minds over. poolside just dropped laguna s 2.1: 118b total parameters, only 8b active per token, a full 1m context window, open weights under a real open license, on huggingface today. look at the chart. it lands at 71 on terminal-bench at 118b, sitting above deepseek v4 pro max at a trillion params, above inkling at 1.5 trillion, above nemotron 3 ultra. it's beating models ten times its size and losing only to kimi k3, which is 24 times bigger. that's the efficiency frontier, up and to the left, exactly where you want a model to sit. but here's the part that made me sit up: it runs on a single dgx spark. and this is what nobody's saying loud enough. the dgx spark is the moe king. a dense 118b would crawl on it, the bandwidth chokes reading every weight each token. a moe with 8b active only ever reads 8b, so the spark's 128 gigs holds the whole model while generation stays fast. big brain, light footprint, the exact shape the spark was built to run. open, frontier competitive, moe efficient, and it fits on a box on your desk. that's the whole thesis in one release: you don't need a datacenter, you need the right architecture on the right hardware. go grab the link below, weights are up.
1
17
1,508
Pengming Wang retweeted
We’re also releasing Laguna S 2.1 Base today, after extending its native context length to 256k tokens, for research purposes. If you’d like access, just reach out to us! Metrics are below.
2
5
66
4,309
Pengming Wang retweeted
new local coding agent model dropped. poolside released laguna s 2.1: 118B MoE, 8B active, 1M context, open weights, openmdw license 70.2% terminal-bench 2.1 78.5% swe-bench multilingual — level with claude sonnet 5 59.4% swe-bench pro runs on a single DGX Spark. up to 24h autonomous runs. here's a fun game I mde with the prompt: "Create a simple video game like Doom but with scary devils. Use JavaScript in an embedded file and save game.html"
6
6
58
7,490
Pengming Wang retweeted
How did we get to 60 days? 1⃣ Compute - we do have quite a bit (...but it is never enough, and meetings around what ablations to run were the most heated ones) 2⃣ Team - people are really smart and hard-working (quite a few weekends these last days) 3⃣The infra - there is a reason we talk about the model factory, it makes iterating much faster (and more pleasant!)
Today we are releasing Laguna S 2.1. At 118B total parameters, with 8B active per token, it does the work of models several times its size on agentic coding. It is remarkably persistent across long-horizon tasks. And it is small enough to run on a single NVIDIA DGX Spark. It is far more capable than anything we have created before, and I think it redefines what a model in its weight class can do. Laguna S 2.1 is an important model for Poolside. What it represents is even more important. If, five years ago, I had read a book that said that by 2030 everything economically valuable, scientifically interesting, and personally meaningful would be built on intelligence contracted from three or four companies, I would have called it dystopian science fiction. We are at a fork in the road of what kind of world we can have. I believe intelligence should and will become a commodity. The question is whether that intelligence comes from three companies, or from many people who can build it, own it, and shape it. The open ecosystem will not win by being the best in its own category. No one cares who is king of the open-source kingdom. People want the best intelligence for the task they are trying to do, with the right balance of quality, speed, cost, and control. If we want a different future, open models have to be on par with, or better than, their closed equivalents. Laguna S 2.1 is a meaningful step in that direction: capable enough to compete far above its weight class, efficient enough to run on hardware you can own, and open-weight so anyone can build on it. Open-weighting our models is the contribution we can make today toward a world where intelligence can be built and owned by many. And we will keep doing it. I am very proud of this team’s work. A big shout out to everyone at Poolside who made this possible, from infrastructure and data to architecture, pretraining, post-training, evaluations, and inference. Laguna S 2.1 is available today under the OpenMDW-1.1 license, with weights on Hugging Face and access through OpenRouter and our API. We are building toward a future where the most capable intelligence in the world can be owned and shaped by anyone. Laguna S 2.1 is one step. We are going to keep building until that future exists. poolside.ai/blog/introducing…
2
2
55
2,858
Pengming Wang retweeted
Laguna S 2.1 is, as far as we can measure, the most capable agentic coding model in its weight class. On Terminal-Bench 2.1 it scores 70.2, sitting beside models 5–25x its size and ahead of several of them. And on DeepSWE from @datacurve, the hardest long-horizon benchmark we ran, Laguna S 2.1 scores 40.4, outperforming some open models with more than 1T parameters. For every score we publish today, we're releasing the full trajectory of every trial in the final evaluation set at trajectories.poolside.ai
18
46
494
128,682
Pengming Wang retweeted
Today, we are releasing @poolsideai Laguna S 2.1, a 118B-A8B open-weight (OpenMDW-1.1) model with a 1M-token context window, completely built in-house for long-horizon agentic work. Laguna S 2.1 is a big step-up over our previous models and is resourceful and persistent. These traits make it the strongest coding agent model at this size, and competitive with models many times larger. For instance, we achieve 70.2% on Terminal-Bench 2.1, which is higher than Thinking Machines Inkling (released last week, 975B, 63.8%) and DeepSeek V4 Pro Max (1.6T, 64%), and closing on Gemini 3.6 Flash (closed-weights, released today, 78%). This carries over to other benchmarks such as SWE-Bench Pro (59.4%, beating Gemini 3.6 Flash at 58.7%). We are releasing all eval trajectories to back these numbers up. Benchmarks are one thing, real-world usage another. What I really like about this model is that you can point it at a task and it will happily churn away for hours. @bmuskalla had Laguna S 2.1 build a working HTML/CSS rendering engine from scratch in a single 50-minute session without any human intervention. Of course, the model has limitations: it overthinks, and it works best in our own harness pool. I'm incredibly proud of our team at @poolsideai. Working on this model and seeing it come together across the different training stages was fantastic. Check out the full release post for all the details, the download links and the free inference via OpenRouter. We are looking forward to hearing all feedback: the good, the bad, and the ugly. Now, back to work.
Today we're releasing Laguna S 2.1, our most capable model to date. It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes. Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark. Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface poolside.ai/blog/introducing…
25
51
534
33,833
Pengming Wang retweeted
Today we're releasing Laguna S 2.1, our most capable model to date. It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes. Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark. Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface poolside.ai/blog/introducing…
220
450
3,537
1,421,970