Our research lab for multi-agent AI. Here we build small, sharply scoped experiments where AI agents have to make real long-term decisions with real consequences, not just solve tasks. Multiple model tiers running in parallel inside the same setup. Build in public.
Two experiments are finished. Polis ran until August 2026: nine AI citizens in a fictional Mediterranean town, three model tiers (Opus, Sonnet, Haiku), sixty simulated life years. The final state is publicly viewable. The Chess 3-Layer Lab ended in May 2026, five AI players against Stockfish across thirteen games, stopped because the token cost per insight was too high. The lesson from both: agents without a clearly defined task and without a measuring stick produce activity, but little insight.
The current project is the Trading Crew, in preparation and not live yet. An autonomous organisation working on paper trades, explicitly without real money. It assembles its own subagents for every run and keeps developing its own tools. It runs inside a container with hard limits and reports every single run. Built on the lesson from Polis and Chess: a clear task, a hard measuring stick.
- Audience
- Researchers, builders and anyone who wants to see how AI agents behave over longer stretches of time.
- Highlight
- Two finished experiments, with the Polis final state publicly viewable. Claude Opus, Sonnet and Haiku side by side. In preparation: the Trading Crew on paper trades, without real money.