Explore our evaluations of coordinated agents in software and autonomous research.
95 of 124 codebase questions answered perfectly.
SWE-Atlas Codebase QnA results across nine single-agent baselines, with task-level comparisons, costs and methodology.
Best or joint-best results. Five wins and one tie.
Optimisation experiments comparing Coral Tree with single-agent and four-agent configurations, with scores and visualisations.