Most AI agents today are disposable reasoners.
They solve a task once, then start from scratch the next time. And when a learned workflow breaks, they often fall back to reasoning from scratch again.
I think agents should work very differently:
๐ฅ๐๐๐ซ๐ง โ ๐ซ๐๐ฎ๐ฌ๐ โ ๐๐๐ข๐ฅ โ ๐ซ๐๐ฉ๐๐ข๐ซ โ ๐ข๐ฆ๐ฉ๐ซ๐จ๐ฏ๐
Introducing ๐๐๐ฎ๐ซ๐จ-๐๐ฒ๐ฆ๐๐จ๐ฅ๐ข๐ ๐๐จ๐ฆ๐ฉ๐ฎ๐ญ๐๐ซ ๐๐ฌ๐, where agents turn execution experience into reusable, continually improving policies.
Instead of repeatedly paying a frontier model to re-plan every action, the agent distills experience into a neuro-symbolic program: stable knowledge becomes executable code, while perception and uncertain decisions remain neural.
More importantly, these policies are not static.
When execution fails, the agent diagnoses what went wrong, reasons about what should have happened, and repairs the policy for future runs.
In that sense, the agent becomes ๐ฌ๐๐ฅ๐-๐ก๐๐๐ฅ๐ข๐ง๐ : failures are not just errors to recover from, but opportunities to improve the underlying skill.
As the agent encounters new parameters, states, and edge cases, the same policy can keep evolving. The policy itself becomes a form of ๐๐๐๐ฎ๐ฆ๐ฎ๐ฅ๐๐ญ๐๐ ๐จ๐ฉ๐๐ซ๐๐ญ๐ข๐จ๐ง๐๐ฅ ๐ฆ๐๐ฆ๐จ๐ซ๐ฒ.
To me, this points toward an important direction for continual learning in agents.
Continual learning doesnโt have to mean constantly updating billions of model weights. It can also mean ๐๐จ๐ง๐ญ๐ข๐ง๐ฎ๐๐ฅ๐ฅ๐ฒ ๐๐ฎ๐ข๐ฅ๐๐ข๐ง๐ , ๐ซ๐๐ฎ๐ฌ๐ข๐ง๐ , ๐ซ๐๐ฉ๐๐ข๐ซ๐ข๐ง๐ , ๐๐ง๐ ๐ซ๐๐๐ข๐ง๐ข๐ง๐ ๐๐ฑ๐๐๐ฎ๐ญ๐๐๐ฅ๐ ๐ฌ๐ค๐ข๐ฅ๐ฅ๐ฌ ๐๐ซ๐จ๐ฆ ๐๐ฑ๐ฉ๐๐ซ๐ข๐๐ง๐๐.
The results:
- ๐๐๐๐ ๐ฉ๐๐ฌ๐ฌ^3 ๐ฌ๐ฎ๐๐๐๐ฌ๐ฌ ๐ซ๐๐ญ๐ across all four settings, improving by 3.6โ15.8 ๐ฉ๐จ๐ข๐ง๐ญ๐ฌ
- Up to 217ร ๐ฅ๐จ๐ฐ๐๐ซ ๐๐ฑ๐๐๐ฎ๐ญ๐ข๐จ๐ง ๐๐จ๐ฌ๐ญ
- Up tp 5.1ร ๐ฅ๐จ๐ฐ๐๐ซ ๐ฅ๐๐ญ๐๐ง๐๐ฒ
Long term, useful agents shouldnโt just be able to reason. They should ๐ซ๐๐ฆ๐๐ฆ๐๐๐ซ ๐ฐ๐ก๐๐ญ ๐ญ๐ก๐๐ฒ ๐ฅ๐๐๐ซ๐ง๐๐, ๐ซ๐๐ฎ๐ฌ๐ ๐ฐ๐ก๐๐ญ ๐๐ฅ๐ซ๐๐๐๐ฒ ๐ฐ๐จ๐ซ๐ค๐ฌ, ๐ซ๐๐ฉ๐๐ข๐ซ ๐ฐ๐ก๐๐ญ ๐๐ซ๐๐๐ค๐ฌ, ๐๐ง๐ ๐ข๐ง๐๐ซ๐๐๐ฌ๐ข๐ง๐ ๐ฅ๐ฒ ๐ค๐ง๐จ๐ฐ ๐ฐ๐ก๐๐ญ ๐ญ๐ก๐๐ฒ ๐ง๐จ ๐ฅ๐จ๐ง๐ ๐๐ซ ๐ง๐๐๐ ๐ญ๐จ ๐ซ๐๐๐ฌ๐จ๐ง ๐๐๐จ๐ฎ๐ญ.
Very proud of the team for pushing toward this vision.
Three years into making computer use agents, I kept asking myself:
Why hasn't CUA reached super intelligence yet?
@sai_borg (Sai), which happens to sound like SI, has invented a new way to build computer use agents: not just more reliable, but also 99% cheaper
The answer lies in how humans function: build muscle memory to optimize for efficient output over time
We released a paper today on this new paradigm, called neuro-symbolic
Give it a read here:
simular.ai/articles/real-worโฆ