The training platform for agents

London & San Francisco
Pinned Tweet
Introducing Overmind. The training platform for agents. Overmind trained models across biomedical, legal, and engineering tasks beating the frontier and achieving: • 7x better accuracy • 20x cheaper usage • 28x less hallucinations Train your own: github.com/overmind-core/ove…
145
42
376
977,229
Who’s building this?
1
4
239
The future belongs to those that build their own intelligence.
Replying to @OvermindLab
Everyone’s building on the same models. Time to turn your team into a frontier lab: ▸ Star the repo: github.com/overmind-core/ove… ▸ Run the cloud version: console.overmindlab.ai ▸ Read more: overmindlab.ai/research/when…
2
6
319
Overmind retweeted
“...big is unnecessary” — Jensen Huang, pitching Overmind
9
2
24
1,018
Overmind retweeted
agent builders are about to realize that production traces are gold. the agents that log failures, evaluate properly, and turn real usage into training data will compound way faster than prompt only wrappers. pretty interesting direction.
Introducing Overmind. The training platform for agents. Overmind trained models across biomedical, legal, and engineering tasks beating the frontier and achieving: • 7x better accuracy • 20x cheaper usage • 28x less hallucinations Train your own: github.com/overmind-core/ove…
2
4
51
“...big is unnecessary” — Jensen Huang, pitching Overmind
9
2
24
1,018
NVIDIA Research making their case: developer.nvidia.com/blog/ho… Overmind Research making ours: overmindlab.ai/research/when…
2
6
143
Introducing Overmind. The training platform for agents. Overmind trained models across biomedical, legal, and engineering tasks beating the frontier and achieving: • 7x better accuracy • 20x cheaper usage • 28x less hallucinations Train your own: github.com/overmind-core/ove…
145
42
376
977,229
Everyone’s building on the same models. Time to turn your team into a frontier lab: ▸ Star the repo: github.com/overmind-core/ove… ▸ Run the cloud version: console.overmindlab.ai ▸ Read more: overmindlab.ai/research/when…
20
2,990
Overmind retweeted
Keep your tracing stack. Train from it. Overmind now imports traces from @langfuse, @LangChain (LangSmith), @braintrust and Galileo. Connect a project and it backfills the history, then keeps syncing. Those traces become eval sets and training data for a model you own.
15
3
29
1,014
Overmind retweeted
The Token Lords have given us such bounty! Honestly the pace of innovation at @OpenAI is astounding.
1
7
237
Keep your tracing stack. Train from it. Overmind now imports traces from @langfuse, @LangChain (LangSmith), @braintrust and Galileo. Connect a project and it backfills the history, then keeps syncing. Those traces become eval sets and training data for a model you own.
15
3
29
1,014
From Galileo: traces, spans and metrics, self-hosted included.
1
5
30
Add a source from the terminal: overmind connector add langfuse Swap in langsmith, braintrust or galileo, or use Observability → Integrations. Then pick the project and sync. Credentials are encrypted and only used to read. docs.overmindlab.ai/core/obs… More on Thursday.
5
34
Overmind retweeted
What counts as a good run for @browser_use? Its own code already spells it out. Overmind builds a context graph of an agent from its codebase and maps its tasks and trajectories. It then generates evaluators from the rules in that code.
1
1
8
436
Overmind retweeted
Fine-tuning a smaller open-weight model isn't a compromise. For specialised tasks it can be the better way to build.
1
1
5
233
Overmind retweeted
Fine-tuned specialist vs frontier flagship: • On contracts, it quoted clauses verbatim 7x as often, at a twentieth of the cost • On BioRED, a fine-tuned Qwen 9B was right 4x as often • On NASA ASRS, a 12B model tied on meaning and won on wording Bigger wasn't better. Think Smaller.
Replying to @TylerEdwas
Across legal, biomedical and aviation benchmarks, our fine-tuned models were 7x more accurate at quoting clauses word for word and 20x cheaper on the legal task. How we did it: overmindlab.ai/research/when…
10
2
17
444