Founder & CEO, Legatus AI 🏛️ The control plane for autonomous software development—from objective to production.

Nashville, TN.
The benchmark I care about for coding agents isn’t “can it generate the code?” It’s: can the system take a messy objective through implementation, testing, review and verification—and leave the repo in a state you’d actually maintain? The demo is generation. The product is delivery.
2
1
41
If an agent can act without waiting for you, it needs a way to prove what it did without asking you to trust its own story. That’s the part of autonomy I think gets underestimated. Execution gets cheaper. Verification becomes more valuable.
1
1
11
One agent making a bad assumption is a bug. Ten agents inheriting the same bad assumption is a systems failure. That’s why scaling agent count without scaling coordination and verification can make the system worse, not better. Autonomy amplifies architecture—good and bad.
1
26
Agent memory isn’t chat history. For autonomous software work, the durable state is decisions, requirements, evidence, failures, approvals and why the plan changed. That state can’t disappear because a context window rolled over. Memory is infrastructure.
1
2
62
I think prompts are becoming the wrong abstraction for serious autonomous software work. A prompt is an instruction. An objective is a destination. Give the system an objective, constraints and verification criteria—then let it plan the path while proving the critical transitions. That’s closer to how I’m designing Legatus.
1
64
Independent verification changes how you build the whole agent system. If completion has to be proven, you stop optimizing for impressive output and start designing around evidence: What was required? What changed? What tested it? Who—or what—verified it? That’s a much healthier definition of “done.”
93
Running more agents in parallel sounds like free speed. It isn’t. You also multiply conflicting assumptions, duplicated work and state synchronization problems. A lesson I keep coming back to while building Legatus: parallelism only helps when decomposition and reconciliation are first-class parts of the system.
62
I don’t want Legatus coupled to whichever model is #1 this month. Model leadership changes too fast. The durable architecture is one where models are replaceable workers: route each job to the right capability without rebuilding the system around a new leaderboard winner. Models change. The control layer should survive.
38
Software delivery isn’t a straight line. It’s a graph. Dependencies. Decisions. Tests. Reviews. Exceptions. That changes how I think about autonomous development: the system has to preserve the work graph, not just the latest conversation with an agent. The prompt is temporary. The graph is durable.
33
A lot of AI failures look like model failures but are really context failures. Wrong file. Old assumption. Missing constraint. Lost decision. One thing building Legatus keeps reinforcing for me: context isn’t a prompt problem. It’s a systems problem. The agent needs the right state at the right moment.
28
Jev is a bigger signal than another fast model. We’ve been using generative LLMs for control flow: route, retry, escalate, verify. That’s backwards. Generation and judgment are splitting into different primitives. The agent stack is becoming a system of specialized models.
1
56
One design decision I’m increasingly convinced of while building Legatus: Keep AI out of the critical execution path. Agents can reason, plan and propose. But state changes, limits and irreversible actions should pass through deterministic controls. Probabilistic intelligence. Deterministic authority.
1
39
There’s a difference between an AI agent completing a task and a system delivering an outcome. Tasks are local. Outcomes are end-to-end. The future belongs to systems that can own the whole path from intent to verified result.
1
50
The strongest AI systems won’t be the ones that never fail. They’ll be the ones that fail visibly, recover intelligently, and preserve enough context to keep moving. Resilience matters more than perfect demos.
1
35
When AI can write thousands of lines of code before lunch, code generation stops being the bottleneck. Decision quality becomes the bottleneck. What should be built? What should change? What should be rejected? What is actually safe to ship?
29
The interesting part of agentic software isn’t getting one agent to do something impressive. It’s getting a system of agents to behave predictably over time. That means memory, state, permissions, handoffs, verification, and recovery.
30
One thing Komodor’s new agentic ops platform gets right: agent changes shouldn’t go straight to production. Golden scenarios. Shadow runs. Promotion only after evidence. We already know how to ship risky software safely. Agent systems need the same release discipline.
17
I think the AI coding market is still focused too heavily on generation. Generation is becoming abundant. Trustworthy execution is scarce. Knowing that the right thing was built, for the right reason, without breaking everything else—that’s the harder problem.
13
A lot of “autonomous” AI still requires a human sitting there watching it. That’s not autonomy. Autonomy means the system can execute, check itself, recover from failure, and know when it actually needs you.
13
Gary Somerhalder retweeted
skills compiling into executable graphs is a sharp architecture split capability selection in front of the workflow is the layer that matters đź’»
1
1
15