The most tempting idea when building with AI is the one good agent. You describe the job properly, give it tools, and it does the rest.
That works surprisingly far and fails at one particular spot: quality-checking its own work.
An agent can miss errors when reviewing its own text because it carries the same assumptions forward. That does not mean self-review finds nothing: Self-Refine demonstrated improvements with feedback from the same model. A separate review assignment, explicit criteria and checking whether review actually improves the result are what matter.
Roles Instead Of All-Rounders#
Crew is our tool for the other route: one job, several agents with different roles.
A typical split looks like this. One gathers what exists on the topic. One drafts. One hunts for weak spots and is allowed to reject. At the end one merges the results.
The reviewer is the most important, and its assignment has to be phrased differently than usual. Not "review the draft" but "find what is wrong with it". The difference sounds like hair-splitting and is the entire effect. The first phrasing gets agreement with minor notes. The second gets findings.
What You Can Get Wrong#
The most common mistake is building too many roles.
Five agents all contributing a bit produce five opinions and no decision. What comes out at the end is a text that has accommodated every note and therefore claims nothing.
For our tasks we prefer starting with a few clear roles. Test whether three outperform six on actual results. A final decision role helps avoid contradictory approvals.
The same brief and context can encourage shared blind spots. A reviewer with separate criteria or additional evidence can help; check its actual effect.
Where We Use It#
Anywhere something comes out that somebody will read. Research that leads to a decision. Texts that get published. Analyses from which we recommend something to a client.
Not where a task is unambiguous. Fetch a number, convert a file, check a status: a single agent is right for that, and three would be waste.
The rule of thumb that has settled here: as soon as a result contains an opinion, it needs a second party attacking it.
Why It Is A Tool Of Its Own#
You can build all of this by hand. We built it by hand for a long time, freshly for every use case.
What is missing when you do it by hand is the trail. Who said what, in which order, what did the reviewer object to, was it taken into account. Without that trail, the next time a result is poor, you do not know where it went wrong.
Crew keeps that trail. That is the actual reason for a tool rather than a collection of scripts.
The Honest Part#
Several agents cost more than one. More time, more compute, more places where something can get stuck.
The maths only works out if the result genuinely gets better, and that is not automatic. Given a badly framed job, three agents produce three badly aimed contributions instead of one.
So the setup does not replace thinking about what should come out. It only makes sure nobody signs off their own work.
