Coral against the baselines
Coral AI Labs All research briefs ↗ · September 2026

SWE-Atlas Codebase QnA · Coral evaluation

Coral against the baselines

Inside one question

What Coral stood up for each question: medians over the 120 fully accounted runs. Every baseline is one agent working alone.

Codexcoordinator6 investigators4 curators369 file agents, each owning one file
RoleWhat it doesAgentsInference callsCalls per agent
Codex (outer agent)Asks the question once, as the developer's agent1181181
CoordinatorOwns the question; will not publish while evidence is outstanding1285283
InvestigatorsEach chases one line of inquiry through the code6770150
CuratorsWeigh and file every piece of evidence4704176
File agentsEach owns one file of the repository3696802.1

Per question: 379 agents made inferences (90 to 449), about 2,764 inference calls in 36 minutes. 79% of file agents make exactly one call. File agents start on Ling 3.0 Flash and a median of 39 per question escalate to DeepSeek V4.1 Flash; the other roles run on DeepSeek V4.1 Flash.

Rubrics

Perfect-score rate

Share of questions where the answer earned every rubric point.

Who is perfect where the other is not

Question by question against each baseline. Questions both answered perfectly, or neither did, are counted on the right.

Only Coral perfectOnly the baseline perfect

Every question, easiest to hardest

One column per question, sorted by how many of the nine baselines answered it perfectly. A filled cell is a perfect answer. Hover a column for the grades.

Coral perfectCompared baseline perfectOther baseline perfectNot perfectBlocked by OpenAI's cyber filter, scored 0Opus 5.5 fell back to Opus 4.8 after a cyber refusal

By repository

Accuracy against inference cost

CoralCompared baselineOther baselinesNo condition beats these on both cost and accuracy

Total cost, people included

What a correct answer costs

Tokens are the cheap part. A wrong answer to a hard codebase question arrives with the same confidence as a right one, no test fails, and a developer builds on it. Priced per 100 questions with the developer time each wrong answer costs, the accuracy gap becomes the cost gap.

What one wrong answer costs:

A missed wrong answer is found later, when something breaks. It then costs the manual answer plus unwinding the work built on it.

US software developer median wage of $65.38 (BLS, May 2025) divided by 0.70, the wage share of employer costs.

Some Coral costs were recorded at off-peak or incomplete rates. Raise this to see how far the result holds.

Total cost per 100 questions

Measured inference plus developer time on wrong answers, at the current settings.

InferenceDeveloper time on wrong answersCoral
Fewer wrong answers
Extra inference
Hours saved
Net saving

The input that drives the cost

What one wrong answer costs a developer

A benchmark shows every wrong answer to the person grading it. Real work does not. These questions ask how a live system behaves, and the answer arrives with the same confidence as a right one, so the developer builds on it.

The cost appears later, when something breaks. Someone traces the failure back to the answer, works out the right one by hand, and unwinds the code, tests and decisions built on the wrong one.

Start1 wrong answergiven with the same confidence as a right one
Caught when given
Missed, then built on
Expected cost of one wrong answer
EstimateDefaultReasoning
Wrong answers caught when given25%Codebase answers cannot be checked automatically. A developer spots a wrong one only if they already know the answer, which is why they asked. Scale describes these questions as requiring the app to be set up, run and traced (paper).
Time to check and fix a caught one30 minRe-read the code paths the answer cites and correct it before anything is built on it.
Time to answer the question by hand2 hoursEach question needs the application running and traced. On SWE-bench Verified, annotators rated 90% of its simpler bug fixes at under an hour (Epoch AI, 2025); SWE Atlas is built to be harder. Developers spend about 58% of their working time understanding code (Xia et al., IEEE TSE, 2018).
Rework on what was built on it+50%Code, tests and decisions made on the wrong premise have to be found and undone. In 2026, 52% of developers report more time debugging problems AI introduced (BairesDev, Q3 2026).

Expected cost = share caught × time to fix a caught one + share missed × time to answer by hand × (1 + rework). At the defaults: 25% × 30 min + 75% × 2 h × 1.5 = 2 h 23 min, or $222 at $93.40 an hour.

52%of developers spend more time debugging problems AI introducedBairesDev, Q3 2026
67%spend more time reviewing AI-generated codeBairesDev, Q3 2026
78%of CTOs increased spending on code review and QABairesDev, Q3 2026
66%spend more time fixing "almost right" AI codeStack Overflow, 2025

Method

How this was measured, and what could change it

The measurement

  • 124 questions from Scale AI's SWE-Atlas Codebase QnA benchmark, drawn from 11 open-source repositories. An answer is perfect when it earns every rubric point.
  • Every condition has a graded result for all 124 questions. Scores in this dashboard come from the GPT-5.5 judge, from Coral's results table of 29 September 2026.
  • The official view uses every rubric as published. The second view leaves out 31 rubrics a review found wrong, where what the rubric requires cannot be stated accurately for the question. It is a reporting view, not a regrade.
  • Answers blocked by OpenAI's cyber filter count as results scoring 0. On 13 questions Opus 5.5 fell back to Opus 4.8 after a cyber refusal; those answers are graded as given.
  • Astra xhigh, Astra max and Astra max 872k are later runs in an isolated environment with the benchmark's official task instruction. The Opus 5.5 xhigh run also used the official instruction.
  • Costs are the selected attempt at frozen API-equivalent prices, judges excluded, shown as the median per question. Questions with an incomplete or unavailable cost are left out; Coral has costs for 92 of 124.

What could change the numbers

  • The cost of a wrong answer is an estimate from four stated inputs until it is measured. The break-even does not depend on it: it states the minutes above which Coral pays off, and every realistic estimate sits far above it.
  • Four studies would turn the estimate into data: expert-timed tasks, as METR does; tasks with a real price, as in SWE-Lancer; time until a known wrong answer is discovered in normal work; and time to fix recorded by coral-code in real teams.
  • DeepSeek-based costs use assumed off-peak rates, and Coral's cost median covers 92 of 124 questions. Both may understate Coral's cost; the stress test shows the headroom.
  • These are Coral's own runs of every condition. Full traces and reproduction materials are available by arrangement. Request reproduction access.
  • Repositories hold 6 to 26 questions each, so single-repository results are indicative rather than precise.