Tech
Multi-Agent Systems: 4 Tests for When One Agent Beats Five
A team at Anthropic built a research system where one lead agent hands work to several subagents running in parallel. On their internal research eval it beat a single agent by 90.2%.
In the same write-up they said it burns about 15 times the tokens of a plain chat, and that token usage by itself explained 80% of the variance in performance on the BrowseComp benchmark.
Put those two numbers next to each other and the headline changes. Five agents did not win because five heads think better ...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to