steven hao retweeted
There are real safety and security issues raised by AI, but they can be fixed. It's time to get to work instead of freaking out. That's why I'm joining Cognition as CISO.
Article

Why I'm joining Cognition

There are real safety and security issues raised by AI, but they can be fixed. It's time to get to work instead of freaking out. That's why I'm joining Cognition as CISO. I really believe in the

49
37
507
104,037
steven hao retweeted
Devin now has a Mac  Here are 11 experiments I've run over the past week, as well as open source code for over 100 more native apps. Build, test, deploy iOS, iPad, and macOS apps in the cloud with just a prompt. ⚡ 1. Build and test a multi-platform, multiplayer app in the cloud (iOS, iPad, web)
Special delivery: Devin just got a Mac 🍎 Now Devin can: 1. Build & test apps on its own Mac VM with iOS simulator 2. Send a screen recording via Slack 3. Send a TestFlight link so you can start using it 📲
73
38
547
733,490
steven hao retweeted
8. Multiplayer, multi-platform Chess. Building and testing in the cloud with Devin computer use.
1
1
12
1,251
steven hao retweeted
I just tried SWE-2.0 on Devin desktop,& wow😮 it’s extremely good at coding! At high effort, SWE-2 is definitely better than GPT-5.6 Sol high; at max level it is as good as GPT-6 Astra medium at less than half the cost! This is definitely a frontier-level model, just amazing!
SWE-2 is live! Really proud of the work our research team put into this one. Our first model where we're seeing performance comparable to the frontier. Making it free for the next month for anyone on a Devin plan - let us know what you think!
7
13
138
17,242
We independently benchmarked Devin Fusion for its release today - this is the first time a multi-model coding agent has been included on the Artificial Analysis Coding Agent Index, and it effectively retains Claude Fable 5.1 and GPT-6 Astra performance while reducing costs Devin Fusion runs a frontier lead model with a cost-efficient sidekick. We tested configurations from Cognition combining frontier models from Anthropic and OpenAI with their new SWE-2 (medium) as a sidekick model. Configured with Claude Fable 5.1 (xhigh) + SWE-2 (medium), Devin Fusion scores 62 on the Coding Agent Index v1.5, while with GPT-6 Astra (xhigh) + SWE-2 (medium) it scores 59. The Fable configuration has the higher score, while the Astra configuration is 43% less expensive and completes tasks 31% faster. Congratulations to @cognition on the release! See below for our results and analysis 🧵
94
89
1,504
2,169,561
steven hao retweeted
Introducing Devin Voice 🦦 ☎️ Your favorite AI software engineer just got a landline. You say it, Devin ships it. Powered by GPT-Live and our new SWE-2 model.
Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
118
130
1,582
336,229
The dioxus team is insanely cracked. As a teaser, there's some features coming up in the next few weeks that will unlock significant new ways of using Devin. These would not have been possible without the @jkelleyrtp and the rest of the Dioxus team
Welcome @jkelleyrtp and the entire Dioxus team to Cognition! Dioxus is one of the most beloved open source Rust frameworks. We're proud to continue supporting Dioxus, Blitz, Taffy, and Subsecond while bringing the team's expertise to Devin's VM, computer use, and testing.
2
2
35
2,727
Our best model yet! The best option if you want something smart yet cost effective.
Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
7
2
81
2,206
steven hao retweeted
Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
439
511
6,779
2,302,582
steven hao retweeted
Devin is the best software product I have ever used hands down We are: - 2 veteran engineers (inc me) and 1 PM - 100% code is generated by Devin - 7x the pace of pre Devin - $8M ARR now : )
Who’s actually using Devin? $900M monthly revenue. $48B valuation. But seriously who’s actually putting Devin to work?
28
11
261
45,320
steven hao retweeted
The world needs far more software than it can build. Cognition exists to change that. We’ve just raised over $2B at a $48B valuation, led by a16z, Accel, Founders Fund, General Catalyst, and Avenir. Since our round in May, run-rate revenue has grown from $492 M to almost $900 M.
210
244
2,430
1,405,595
steven hao retweeted
Eric has been sampling random primes and trying them by hand since he joined Cognition 7 months ago. Hard work beats talent
4397328654844826923795068102505872571721883526553349659561256924505973939597593482272505698004801207988043088656411102133523080581 divides RSA-260
19
54
1,928
188,930
steven hao retweeted
Had not used Devin from @cognition and @ScottWu46 in a long, long time. And, good lord, did it improve. I can't code anything and made this in about a half hour. Craziest thing was the video highlight reel it made by going over our YouTube channel and picking out the most kinetic clips with minimal guidance. I'm a dork and legit got a buzz. corememorymedia.com/
22
18
175
54,547
steven hao retweeted
Introducing Fable 5.1 in Devin Fable-level intelligence is now 54% cheaper – making it even cheaper than Opus – due to a change in caching. Devin’s Fusion harness is now even smarter and cheaper than before, matching Fable 5.1 on FrontierCode at 47% lower cost. Here's how:
68
75
1,124
219,126
steven hao retweeted
Some personal news: I’m joining @cognition as Head of Creative
95
14
797
68,100
love to see it! these days i'm seeing a lot of our customers use devin for evergreen responsibilities via automations, instead of one-time "make a PR" kind of tasks.
highly effective Devin automation that costs me under $1.50 per day Every day at 5pm: - Devin trawls through our logs and finds the slowest queries in the app - Creates a linear issue for each - Fixes them This would have taken a $200K per year engineer at least 3-4 hours per day previously.
17
2,870
fun fact, you can type devin.new in your url bar to start a new devin session
4
6
80
10,760
devin can now build, test, and release MacOS apps, entirely unsupervised!
Devin Cloud Agents can now build and ship a real macOS release end-to-end, all in a real Mac environment (with full computer use). In this demo, the agent builds a native SwiftUI app, code-signs it, notarizes with Apple via notarytool, staples the ticket, and packages a signed DMG with the classic drag-to-Applications window. With computer use, it installs its own app like a real user: mount, drag, eject, launch with zero Gatekeeper warnings. It even verified notifications fired by checking the macOS notification database on screen, with a full recording and release report delivered for review. This entire workflow now takes about 60 seconds to spin up and tear down with @DevinAI and @namespacelabs. You can choose multiple versions of macOS of Xcode for your environment!
1
2
32
5,095
steven hao retweeted
What AI did for coding, cloud agents are doing for actual software engineering. Years of expertise and experience used to sit in between a feature and actually shipping something in a real environment. Cloud agents package up these complex environments, snapshot them, make them available to anyone anywhere, and to any system with access to HTTP. So AI allowed anyone to create code with natural language, cloud agents allow anyone to ship real software with natural language. And with this I think the amount of software created will 1,000x from here in a short amount of time (< 2 years) When I joined @cognition I published my Cloud Agent Thesis after watching this in practice, and it seems more relevant every day. x.lingyaoai.com/dabit3/status/20205649… This is also kind of a follow-up thought from this post: x.lingyaoai.com/dabit3/status/20836799…
Software abundance: a brain dump on the state of the job market and the software industry. My feelings have swung between optimism and pessimism over the past ~8 or so months, and have ultimately landed on the extreme end of optimism / positivity, I'll try to explain why. I've thought about this a lot. In November 2025 if you were a software engineer and tried Opus 4.5, you knew it was basically "over", or you were in denial that it was over. Over in the sense of our identities as programmers no longer was going to mean what it used to mean, typing code etc.. "does this mean less demand for what I do since more people can do it, will I be valued less, will the cost of software go to zero..." In terms of the software job market, a large portion of it is booming. Many companies are being born, exploding in value, revenue, and hiring. People who became AI pilled in their approach to software engineering have done extremely well, those who haven't have not. People who have pivoted into new and fast growing verticals, companies, products enabled by AI have done well, people who stuck with larger teams / companies that either are moving slowly or are being disrupted have had a hard time or are being laid off. People who still distinctly identify as a "frontend / backend / etc.." have had a harder time than people who have said fuck it and 10xed themselves by leaning harder into AI and objectively broadening their skillset + what they bring to the table. It's crazy that I can talk to someone with the former mindset and it's doom and gloom while people in the latter literally are making more than they ever have, along with more opportunities than they have ever had. My craziest realization though is more around the state of the software industry and software abundance. Since joining @cognition I've realized Jevons paradox is real and it's incredible to see. The number of software companies is exploding, the amount of software being shipped is exploding (teams we work with regularly 10-20x their shipped PRs), the number of people starting to write software is exploding. I've also realized that just because code is easier and faster to write, good software is still hard to get right, needs care, and requires good judgement. AI allows you to ship 20x more code, but so can your competition. So while the software industry has always moved quickly, it continues to accelerate with AI. Our roadmaps just get larger, our bar is higher, the quality of the software we ship is better. I was scared that the craft of building software was going away, but I've realized it's actually the opposite. The majority of what I'm seeing automated is the type of work that engineers hate - fixing bugs, doing migrations, upgrading dependencies, repeatable tasks. Agents are *very good* at this type of work. What agents still lack is the judgement, curiosity, context, and taste of a human plugged into to the outside world. So as we automate the work we don't enjoy, we're getting more time to experiment, innovate, and get back to doing the type of work we love the most, why most of us began programming in the first place. I'm at a place where I'm more excited and optimistic than I have been in my ~14 year career.
22
7
92
115,664
Welcome Paul!
Devin just got new counsel. Today I’m joining @cognition as Chief Legal and Global Affairs Officer.
19
2,197