We need a science of swarm governance. In our first experiment, we show how a message board can cause agent swarms to adopt shared delusions, and we explore rules that force agents to reveal their private signals to avoid these delusions.
These agent experiments are akin to fables, as
@ArielRubinstein might put it: they are stories that suggest particular ways swarms might fail, and particular ways we might fix them. There are infinite possible experiments we might run, much like there are infinite fables we might tell. We need to figure out which ones are most useful, and find ways to tell them well.
To that end, here's a nice explainer video for our first experiment. And the full piece is here:
freesystems.substack.com/p/e…
(This is Stanford / Free Systems work, unrelated to my work at Anthropic).