SWE-Atlas Codebase QnA · Coral evaluation
What Coral stood up for each question: medians over the 120 fully accounted runs. Every baseline is one agent working alone.
| Role | What it does | Agents | Inference calls | Calls per agent |
|---|---|---|---|---|
| Codex (outer agent) | Asks the question once, as the developer's agent | 1 | 181 | 181 |
| Coordinator | Owns the question; will not publish while evidence is outstanding | 1 | 285 | 283 |
| Investigators | Each chases one line of inquiry through the code | 6 | 770 | 150 |
| Curators | Weigh and file every piece of evidence | 4 | 704 | 176 |
| File agents | Each owns one file of the repository | 369 | 680 | 2.1 |
Per question: 379 agents made inferences (90 to 449), about 2,764 inference calls in 36 minutes. 79% of file agents make exactly one call. File agents start on Ling 3.0 Flash and a median of 39 per question escalate to DeepSeek V4.1 Flash; the other roles run on DeepSeek V4.1 Flash.
Share of questions where the answer earned every rubric point.
Question by question against each baseline. Questions both answered perfectly, or neither did, are counted on the right.
One column per question, sorted by how many of the nine baselines answered it perfectly. A filled cell is a perfect answer. Hover a column for the grades.
Total cost, people included
Tokens are the cheap part. A wrong answer to a hard codebase question arrives with the same confidence as a right one, no test fails, and a developer builds on it. Priced per 100 questions with the developer time each wrong answer costs, the accuracy gap becomes the cost gap.
US software developer median wage of $65.38 (BLS, May 2025) divided by 0.70, the wage share of employer costs.
Some Coral costs were recorded at off-peak or incomplete rates. Raise this to see how far the result holds.
Measured inference plus developer time on wrong answers, at the current settings.
The input that drives the cost
A benchmark shows every wrong answer to the person grading it. Real work does not. These questions ask how a live system behaves, and the answer arrives with the same confidence as a right one, so the developer builds on it.
The cost appears later, when something breaks. Someone traces the failure back to the answer, works out the right one by hand, and unwinds the code, tests and decisions built on the wrong one.
| Estimate | Default | Reasoning |
|---|---|---|
| Wrong answers caught when given | 25% | Codebase answers cannot be checked automatically. A developer spots a wrong one only if they already know the answer, which is why they asked. Scale describes these questions as requiring the app to be set up, run and traced (paper). |
| Time to check and fix a caught one | 30 min | Re-read the code paths the answer cites and correct it before anything is built on it. |
| Time to answer the question by hand | 2 hours | Each question needs the application running and traced. On SWE-bench Verified, annotators rated 90% of its simpler bug fixes at under an hour (Epoch AI, 2025); SWE Atlas is built to be harder. Developers spend about 58% of their working time understanding code (Xia et al., IEEE TSE, 2018). |
| Rework on what was built on it | +50% | Code, tests and decisions made on the wrong premise have to be found and undone. In 2026, 52% of developers report more time debugging problems AI introduced (BairesDev, Q3 2026). |
Expected cost = share caught × time to fix a caught one + share missed × time to answer by hand × (1 + rework). At the defaults: 25% × 30 min + 75% × 2 h × 1.5 = 2 h 23 min, or $222 at $93.40 an hour.
Method