The open platform for automating development. Infrastructure to build, measure, and interact with agents across the SDLC

Agents are getting complex to manage: models, skills, automations, permissions, cloud environments, orchestration patterns... So, we built the Terraform for agents: one repository that configures all of your agents as code. The Warp Factories configuration spec defines: - Agents and orchestration: how work gets routed and which models agents use - Automations: the events and schedules that kick off work - Access and environments: repositories, secrets, MCP servers, and runners - Measurement: scorers and benchmarks for understanding and improving performance Because it’s all code, both engineers and agents can propose changes to the factory itself. That makes it possible to test new configurations, benchmark them, and even have your factory propose improvements as PRs.
9
6
73
156,184
GPT 6.1 Sol is now in Warp, ready to use with your Codex / ChatGPT subscription
56
4,460
You can now sign in with ChatGPT from the Warp Terminal and the Warp Agent CLI! Sign in to Warp with your ChatGPT account and use your subscription’s included Work and Codex usage for your agent conversations.
12
15
161
13,615
Sonnet 5.5 is now in Warp and the Warp Agent CLI We ran the numbers, and confirmed it has reached 100% SOTA on the "Write Sonnets about Rust" bench
3
1
28
6,185
Grok 4.7 has landed in Warp and the Warp Agent CLI. Connect your @grok subscription to get started
3
2
25
7,519
Claude Opus 5.5 is now available in the Warp Terminal and the Warp Agent CLI.
1
2
49
5,112
GPT 6 Sol and Luna are now available in the Warp Terminal and the Warp Agent CLI
2
28
3,830
Finally, use these metrics to have agents suggest improvements to your setup automatically. Self-improvement agents run on a schedule and review failing grades to suggest changes to your setup. Here is an example agent skills PR with evidence cited from previous scoring runs:
1
1
883
Then, aggregate scores to view success and failure trends overtime. This helps observe regressions and track improvements as you ship changes to your factory. Here you can see dips in our redundant test quality, with visibility into failed runs for deeper investigation:
1
1
424
There are many different scorers that you can define. These are the ones we'd start with: Compliance: did the agent complete the task per the user’s request? Efficiency: did the agent complete the task efficiently, or did it do a bunch of unnecessary work? Verbosity: did the agent write succinct code with clear code comments, or was the output unnecessarily elaborate? Quality: for a coding task, was the quality of the code good? Did it match expected conventions?
1
1
520
Then, create your scoring agents. These analyze individual conversations and give them a grade. Scorers follow rubrics that you define. Here's a "redundant test" scorer we use internally to judge the quality of tests that our agents write: gist.github.com/bholmesdev/d… 🔖
1
4
836
Start from a record of agent conversations you want to grade. In Warp Factories, all agent runs are tracked across your team with agent traces, artifacts, and code output. You can score all of your agent runs, or define a sampling you want to test (ex. frontend tasks)
1
2
1,198
Introducing Scorers: Agents that grade your agents. Use LLM-as-a-judge to grade past coding agent sessions on: • Quality • Efficiency • Compliance • Or any custom dimension Scores feed into performance measurements and automatic self-improvement for software factories
18
13
191
568,568
Warp now has built-in support for the Grok Build CLI. - Use Warp's rich input for agent prompts, with support for longer pasted prompts and multi-cursor - Use /remote-control to share your agent session to another device - Access the file explorer and code review panels
27
5
162
973,200
Our company’s cost-per-PR dropped from $80 to $30 by switching to GPT 5.6 Sol. We built a benchmark to replay our team's agent runs across model providers. GPT 5.6 Sol got the highest code quality at a 66% lower cost vs. our old default (Claude Opus 5).
9
4
61
7,449
Factory benchmarks let you build your own model bench using your past coding agent runs. - Mirrors your environment, secrets, and MCPs for each run - Scores output on judging criteria you define - Generates a report with model recs using cost vs. quality Here's how it works:
10
4
58
73,043
We used a factory benchmark to reduce our own cost-per-PR by 63%, from $80 per PR down to $30. Here's a full walkthrough of the benchmark: warp.dev/factories/benchmark… 👀
1
7
1,268
You can use results to inform your default models, or define custom model routers to let your agent pick the right model for each task. For example, you may benchmark frontend UI changes in your codebase, and define a routing rule to always send those tasks to Grok 4.6 medium
1
5
691
Then, review your benchmark report. This includes the overall model recommendation + a performance breakdown across tasks. You can see a pareto graph along each scored dimension (ex: cost vs. quality), and how models performed on each task across your scoring criteria
1
5
719
With your benchmark defined, you can customize the models and scoring criteria you want to test. Use built-in scorers like correctness, code quality, and efficiency, or define custom rubrics for metrics your team cares about (Figma mockup alignment, e2e test quality, etc)
1
6
1,135
Benchmarks are built from your team's past agent runs. All agent runs are tracked in Warp Factories, letting you build a representative sample to test against. You can select tasks and define judging criteria yourself, or build a sample agentically using the Warp Factories MCP
1
12
1,555
Introducing Factory Benchmarks: The first model bench generated from your own coding tasks. Measure, test and improve coding agents by replaying past agent runs, and cut cost-per-PR by 63%+ Here’s how it works 🧵
27
12
188
70,278
Claude Fable 5.1 is now available in Warp and the Warp Agent CLI
12
3
62
7,592
Introducing self-improvement loops. The concept is simple: what if agents could improve Skills by reviewing past conversations? Here's the three step loop: - Score conversations from criteria you define - Isolate failures - Generate skill improvements
18
18
276
129,571
We're moving our agents to the cloud to improve quality and manage cost. To do this, we needed a way to configure environments, harnesses, and security permissions as code. This is configuration format we landed on:
3
3
61
44,218
We shipped Warp Factories in early access this week. Want to see how it works? @BHolmesDev will be live on X and YouTube later today to show cloud software factories in action, with room for Q&A
3
2
29
4,895
What's it like to work at Warp? - Product ownership in every role - High agency, figuratively and literally - Tight collaboration, in-person or remote Join us: warp.dev/careers
6
6
83
20,786
Measure factory performance with cost and velocity metrics, accessible from the factory dashboard and via our API and SDK. Plus, configure self-improvement agents to grade performance and automatically improve your factory overtime.
1
16
2,748
All of these agents have access to computer use on Linux and Mac to reproduce issues and prove correctness of changes. The agent can share recordings on pull requests and within conversations, so you never need to pull down code to verify a change. Here's a real example:
1
12
2,520
Bring whatever model or harness is best for your workflow. You can use Warp’s SOTA agent for multi-model access, including open-weight models, or you can directly run Claude Code or Codex as the harness.
1
12
2,161
Work enters your factories from the tools your team already uses: - Communication tools like Slack or Teams - Task trackers like Linear or Jira - Source code forges like Github or Gitlab - Terminals, IDEs and other local coding agents via the Warp Factory MCP
1
15
3,351
Factories are defined as code and built to integrate with the tools you already use. Think Terraform for agent configuration.
3
29
4,681
Introducing Warp Factories: open, flexible infrastructure for building cloud software factories. - Configure your factory as code - Use any model and any harness - Measure quality with evals and benchmarks on your own data - Built-in self-improvement and memory
40
41
443
302,664
Tomorrow...
8
10
193
17,660
We wanted to build the perfect coding agent for SSH. No extra CLIs to install on the remote machine. Just run `! ssh` and work alongside the agent. Warp eng Kevin Yang shows how it works:
4
9
113
15,611
Agents like Claude Code can run terminal commands for you. But what if you could invite agents like CC into your already running terminal sessions? Ask for agent help in vim, in SQL REPLs, in other TUIs? That's what the Warp Agent CLI does. Eng Kevin Yang shows how it works:
12
10
163
22,325
Grok 4.6 is now available in Warp and the Warp Agent CLI. Run /connect-grok in Warp to sign in with your X Premium or SuperGrok subscription to get started right away.
49
146
942
373,351
Replying to @xsetxz @theo

ALT Moone Boy Waiting GIF by HULU

1
67
Run /connect-grok in the Warp Agent CLI to use your X Premium subscription with Warp!
23
17
214
1,169,315
The new Warp Agent CLI is yours to shape. Here's Moira, one of our engineers, walking through customization! Everything lives in a settings.toml file, and the agent can edit it for you with our bundled skills. Add custom themes, animations, and a status line that you arrange however you want.
6
1
57
6,020
We've reimagined agent orchestration UX with our new Warp Agent CLI, making it easier than ever to see your subagents. Here's Moira, one of our engineers, walking through some of our features! Configure the model, harness, and environment your subagents run with all with a couple shortcuts
9
14
167
53,369
...Okay but can your agent quit vim?
3
2
56
7,750
Introducing custom model routers. Set your own routing rules in plain english: - "Send simple requests to minimax" - "Send planning tasks to Claude" - "Send bug fixes to Kimi k3" And every prompt gets auto-routed to the right model for the job. Here's how to make one:
15
6
97
201,673
You can also configure your own routing logic. Describe your routing rules in plain english and select your preferred models. BYOK and Grok subscriptions are supported too
1
3
929
The CLI also supports multi-agent orchestration with a UI to monitor all of your agents. Pick the model, the environment (local or cloud), and even the harness (claude code and codex subagents are supported)
2
3
270
The agent can also help you with interactive commands like vim sessions or SQL REPLs. Just tag in an agent to let it take control. It's like computer use for the terminal
1
6
587
It's main advantage over harnesses like Claude Code is full shell access. At it's simplest, that means you can manually cd to different directories in the same session, with @ file context and git status updating as you go
1
10
977
Happy National Intern Day to Warp's engineering + design interns @jaidenratti @JerryDizs & John! It's always impressive how quickly interns go from opening their first PR to owning meaningful projects. Blink once and they're reviewing code, fixing bugs, and shipping features at #warpspeed alongside the rest of the team. Thanks for everything you've built this summer, and for making the office a more fun place to be! 😄
6
3
44
12,122
Kimi K3 is now available in Warp. It's the best open-source model we've tested with 13% better task completion than any other OSS model.
3
1
85
10,383
Try it in Warp:
8
1,797
Claude Opus 5 is live in Warp, and we're seeing a 43% reduction in cost per task at nearly the same quality as Fable 5 in our agent evals.
2
4
36
6,909
A contributor has helped us recently ship one of the most-requested features in Warp's history: support for OSC 8 hyperlinks! This is an issue that has been open for 5 years since issue #176. 🧵↓
5
2
44
14,149
Curious what software factories are? Or how you might move agents to the cloud? @captainsafia will be LIVE tomorrow on our YouTube and X to talk all about it!
1
1
22
4,845
Is the software industry cooked? @steipete and @zachlloydtweets have thoughts:
6
5
70
10,111
Warp 🤝 Sentry Use the @sentry MCP to go from a Seer root cause analysis to a PR— without leaving Warp
1
4
47
7,964
Here's how AI software factories work in under a minute 👇
2
6
29
5,544
Finally, you can drag tabs into and out of the terminal window 🙌
9
7
166
10,042
You can now use the full GPT 5.6 fleet in Warp: Terra, Luma, and Sol. It's... not great at puns, but wow is it good at coding
3
2
67
6,371
What is a software factory? @steipete explains, and shows how they (accidentally) built one to maintain OpenClaw:
2
3
41
9,526
Warp now supports Grok 4.5. You can use it by plugging in your X Premium subscription. It's smart, and token output is fast. This is real-time:
9
7
210
362,762
Stop using Claude Fable 5 or "ultra code" for every task! @steipete and @DynamicWebPaige suggest orchestration: - SOTA models to break down workflows - smaller models for execution ...okay unless you're buying an airfryer
7
9
105
47,661
It was great meeting everyone at our Crafting Software Factories event last night. Nearly 2k RSVPs and 300 attendees! So grateful for the excitement and discussion. Until next time SF 👋
4
4
47
6,829
Fable 5 is so back [in Warp]
Warp now supports Claude Fable 5. Fable 5 has Mythos-level performance and is highly capable of /goal-oriented, unsupervised tasks. Ready to embed into your next loop 🔁
2
4
69
8,097
The Warp Agent now supports Claude Sonnet 5. We're seeing significant gains in agentic coding and professional work, all at Sonnet pricing. A good default to set for tasks around your terminal!
1
4
65
5,251