1. AgentRadio
We tested 124 SWE-Atlas codebase Q&A tasks across 11 codebases and four languages. Task success required passing all checks.
All configurations used Claude Code. The Opus full configuration cost $19.45 per task; one agent cost $2.96. Best-of-six independent attempts scored 37.9% at $17.76. We compared these coordination configurations using Claude Code, with near-matched spending in the DeepSeek comparison.
Read the AgentRadio paper ↗