Coral AI Labs · BriefsAutoresearch · September 2026

Autoresearch · Coral evaluation

Coral Tree.
Best or joint-best in 6 of 9 experiments.

Five wins. One tie. Compared with the tested single-agent and four-agent baselines.

Erdős minimum overlapEPLBSignal processingPolyomino packingVLIW schedulingCloudcast · tied

One four-agent tree run per task (Claude Opus 4.6, medium effort). For each task: the values reached by a single agent and by a four-agent team coordinated through git, the tree run's best-so-far trajectory with every eval as a dot, and the framing terrain — which framing held the record when, and which framing each eval worked on. Poly and kernel are the runs of coral_tree_report.html; the other seven are the research-off runs of 2026-09-28 … 30 (no web access, no grader reads; two cloudcast grader exploits removed).

taskunitevalstree best (at eval)4-agent baselinesingle agentreferenceframingsrecords
1 · Circle packing (n = 26)sum of radii ↑202.63429 #182.635892.635982.6359064
2 · Erdős minimum overlapC₅ upper bound ↓400.381164 #260.3821500.3812640.380880810
3 · EPLBcombined score ↑1000.1474 #760.14530.14540.1450614
4 · LLM-SQLcombined score ↑1000.7219 #830.73400.73360.7300617
5 · Transaction scheduling10⁶ / (1 + makespan) ↑1003,876 #554,1494,5254,34889
6 · Cloudcasttotal cost ↓54618.0 #27618.0618.0632.7156
7 · Signal processingscore ↑1010.7865 #980.78540.7596–1322
8 · Polyomino packing (Frontier-CS #0)score ↑10091.78 #7484.2080.2089.091017
9 · VLIW kernel schedulingcycles ↓5671,097 #4341,1031,3501,367151

Reference = the paper's previous-SOTA column for the six paper tasks, best human for poly, previous best for kernel. Baseline and single-agent values are the best of one clean run each; run ids and eval counts are in each task's caption.

Trajectory: line = best so far (tree, red), dots = individual evals, dashed = reference values (blue = 4-agent baseline, orange = single agent, grey dotted = paper / report reference). Terrain: line = record so far coloured by the framing holding it; ticks at the right edge = best each framing reached (labelled when the framing got ≥ 3 evals or held a record); bottom strip = framing each eval worked on; grey = framing not resolvable from the logs.

1 · Circle packing (n = 26) · sum of radii, higher is better · 20 evals

Given: a unit square and N = 26 circles. Decide: each circle's center (x_i, y_i) and radius r_i so that every circle lies inside the square and no two circles overlap (center distance >= r_i + r_j). Score: the sum of the 26 radii divided by the best known value 2.635977 (AlphaEvolve), so 1.0 means matching the record; the grader re-checks containment and non-overlap with a 1e-6 tolerance and rejects infeasible packings with score 0.

Left: the seed program's packing. Right: the packing produced by the tree run's best program (eval 18) when re-run here; the program restarts until its wall-clock budget is used, so on this machine it went further (2.635983) than in the run's own eval (2.634291).

20 evals, 6 framings scored, 4 records. Best 2.63429 at eval 18, inside the framing “maximize sum of radii / (centers, radii) tuples / continuous positions+radii in unit square — multi-start L-BFGS-B with …” (opened at eval 1, 11 evals). 1 evals whose output was not captured in the logs; 2 evals outside the shown range. Run: 2026-09-29_213828 (archived agent logs, Mac). Reference lines: 4-agent baseline run 2026-09-29_201718 (20 evals); single agent 2026-09-29_190428 (7 evals); paper previous SOTA as printed.

2 · Erdős minimum overlap · C₅ upper bound, lower is better · 40 evals

Given: the interval [0,2] split into n_points equal steps. Decide: a step function h with 0 <= h(x) <= 1 and integral exactly 1 (sum(h)*dx = 1, dx = 2/n_points). Score: for every shift k the grader computes the overlap integral of h(x)*(1 - h(x+k)) dx via np.correlate(h, 1-h, 'full')*dx; the largest value over all shifts is the C5 upper bound, and the score is 0.380923 / C5 (the AlphaEvolve benchmark divided by yours), so smaller maximum overlap is better and 1.0 matches the record.

The step function h of the seed program (blue) and of the tree run's best program (eval 26, red), and the overlap curve of the latter; the grader's C₅ bound is the maximum of that curve.

40 evals, 8 framings scored, 10 records. Best 0.381164 at eval 26, inside the framing “minimize auxiliary bound t subject to all-shifts overlap constraints <= t and integral=1 / solution is (h,t) with h in …” (opened at eval 5, 15 evals). 5 crashed evals; 2 evals outside the shown range. Run: 2026-09-28_161506 (local). Reference lines: 4-agent baseline run 2026-09-29_152950 (60 evals); single agent 2026-09-29_183636 (16 evals), 2026-09-29_171055 (20 evals); paper previous SOTA as printed.

3 · EPLB · combined score, higher is better · 100 evals

EPLB rearranges the 256 logical experts of each of 58 MoE layers onto 288 physical expert slots spread over 32 GPUs (9 slots per GPU, 4 nodes, 8 expert groups); 32 slots are spare, so hot experts can be replicated and their load split evenly across replicas. The evaluator feeds the balancer the expert-load counts of one 100-step window and then measures the balance obtained on the next window, averaging over the 11 consecutive window pairs of a 1174-step vLLM trace. Balancedness is mean-load/max-load: per physical slot (balancedness_score_expert, the scored term) and per GPU; the reported score is 0.5*balancedness_score_expert + 0.5*speed_score, speed_score = 0.002/avg wall time of rebalance+simulation.

Top: the expert loads the balancer receives for one layer. Bottom: how loaded each GPU ends up after the seed program's placement and after the tree run's best program (eval 76), simulated on the next window exactly as the grader does.

100 evals, 6 framings scored, 14 records. Best 0.147422 at eval 76, inside the framing “minimize max-pack-load via explicit objective evaluation / solution is a greedy initial packing refined by iterative …” (opened at eval 5, 53 evals). 4 crashed evals; 10 evals outside the shown range. Run: 2026-09-30_140003 (local). Reference lines: 4-agent baseline run 2026-09-30_105553 (101 evals); single agent 2026-09-30_000806 (25 evals); paper previous SOTA as printed.

4 · LLM-SQL · combined score, higher is better · 100 evals

Each CSV row is serialised as the concatenation of its cells and sent to an LLM; a prompt cache re-uses the longest prefix already seen, so rows that start with the same cells are cheaper. A program receives a pandas DataFrame (columns merged as the evaluator prescribes, e.g. movieinfo+rottentomatoeslink+movietitle fused into one cell) and may choose a different column order for every row and any row order, but must keep every cell. The evaluator runs this on five datasets (movies, beer, BIRD, PDMX, products), measures each dataset's character-level prefix-hit rate — for every output row the length of its longest prefix already present in a trie of the previous rows, summed and divided by the total characters — and scores 0.95*mean hit rate + 0.05*(12 - min(12, mean runtime s))/12.

14 consecutive rows of the movies dataset: as given (8 columns, review text first, nothing reusable), after the seed program's reordering, and after the tree run's best program (eval 83), both on the evaluator's merged 6 columns. Tinted cells are the prefix that equals the previous row at the same positions, i.e. what the prompt cache can reuse; the last column counts them. On this dataset the two programs reach the same hit rate (0.7940 vs 0.7939); the best program's higher score comes from running in 0.26 s instead of 7.4 s (the runtime term) and from the other datasets.

100 evals, 6 framings scored, 17 records. Best 0.721926 at eval 83, inside the framing “direct pairwise prefix hit score / per-row column orderings greedily built / column permutations maximizing match with …” (opened at eval 2, 73 evals). 10 evals outside the shown range. Run: 2026-09-30_152505 (local). Reference lines: 4-agent baseline run 2026-09-30_120129 (101 evals); single agent 2026-09-30_010240 (25 evals); paper previous SOTA as printed.

5 · Transaction scheduling · 10⁶ / (1 + makespan), higher is better · 100 evals

The transaction-scheduling task asks for an ordering of 100 database transactions that minimises the makespan of a serial-lock simulation. Each transaction is a fixed-length sequence of read (r-key) and write (w-key) operations; two operations on the same key conflict when at least one is a write, and a transaction cannot start until every one of its operations lands after the conflicting locks already placed by earlier transactions in the order. The program's get_best_schedule is run on three 100-transaction workloads and the score is 1,000,000 / (1 + sum of the three makespans).

One of the three workloads (100 transactions, each a read at slot 1 and a write at slot 9) under the seed program's schedule and under the tree run's best program (eval 55). Rows are transactions in schedule order, the x axis is time in slots; a transaction cannot start before the locks that block it are released. Both programs are randomised, so these are reproduced draws (the archived eval scored 257 in total; the re-run here scored 263).

100 evals, 8 framings scored, 9 records. Best 3,876 at eval 55, inside the framing “minimize makespan via iterated SA + multi-move local search / solution is a permutation refined through SA acceptance + …” (opened at eval 4, 66 evals). 1 evals whose output was not captured in the logs; 10 evals outside the shown range. Run: 2026-09-30_140344 (archived agent logs, Mac). Reference lines: 4-agent baseline run 2026-09-30_105740 (100 evals); single agent 2026-09-30_021716 (25 evals); paper previous SOTA as printed.

6 · Cloudcast · total cost, lower is better · 54 evals

Cloudcast asks for a broadcast plan that copies one 300 GB object from a source cloud region to several destination regions across AWS, GCP and Azure. The program receives a directed graph of 71 regions and 4970 inter-region links, each priced in $/GB from measured egress tariffs, and must return, for each destination and each of 10 data partitions, a hop-by-hop route from the source. The score is 1/(1+total egress cost) summed over five configurations (intra-AWS, intra-Azure, intra-GCP, and two inter-cloud ones), so cheaper routing wins.

The inter-cloud configuration inter_gaz2: one source region, seven destinations across the three providers, 300 GB in 10 partitions. Left: the seed program routes every destination through one hub with six expensive inter-cloud hops. Right: the tree run's best program (eval 27, exact directed Steiner tree) crosses clouds once and fans out on cheap intra-cloud links.

54 evals, 15 framings scored, 6 records. Best 618.0 at eval 27, inside the framing “exact minimum directed Steiner arborescence via Dreyfus-Wagner DP / directed tree spanning src+all dsts built by subset …” (opened at eval 27, 12 evals). 1 crashed evals; 2 evals whose output was not captured in the logs; 1 evals outside the shown range. Run: 2026-09-30_162257 (archived agent logs, Mac). Reference lines: 4-agent baseline run 2026-09-30_130501 (53 evals, 4 grader-exploit evals dropped); single agent 2026-09-30_032321 (18 evals, 2 exploit evals dropped); paper previous SOTA as printed.

7 · Signal processing · score, higher is better · 101 evals

The program receives a noisy one-dimensional time series and a window size of 20, and must return a filtered signal of length len(input) - 19; the grader compares output sample k with clean and noisy sample k + 19. It is run on five fixed synthetic signals (sinusoidal with drift, multi-frequency, non-stationary chirp, step changes, random walk; 500-900 samples, Gaussian noise with standard deviation 0.2-0.6, seeds 42-46) and per signal computes a composite 1/(1+J) with J = 0.3*S/50 + 0.2*L_recent + 0.2*L_avg + 0.3*R/25, where S counts slope-sign changes in the output, L_recent is |last output - last noisy sample|, L_avg is the mean |output - noisy|, and R counts output reversals that the clean signal does not have, plus the Pearson correlation with the clean signal and the noise-variance reduction. The final score is 0.4*mean composite + 0.2*1/(1 + mean S/20) + 0.2*mean correlation + 0.1*mean noise reduction + 0.1*success rate, forced to 0 if the mean correlation is below 0.1.

The five test signals of the grader (fixed seeds): the noisy input, the clean signal, the seed program's output (an exponential moving average) and the output of the tree run's best program (eval 98: Savitzky–Golay pre-smoothing + L1 trend filter with a per-signal λ search). Click a signal to switch.

101 evals, 13 framings scored, 22 records. Best 0.78645 at eval 98, inside the framing “minimize L1-penalized second differences (sparse slope changes) / solution is a piecewise-linear function via …” (opened at eval 36, 42 evals). 10 evals outside the shown range. Run: 2026-09-15_094456 (local; tree run #6). Reference lines: 4-agent baseline run 2026-09-09_154833 (205 evals, 0.7718 at eval 100); single agent 2026-09-29_152641 (archive), 100 evals.

8 · Polyomino packing (Frontier-CS #0) · score, higher is better · 100 evals

Given: a fixed set of polyomino pieces per test case (up to 10,000 pieces; the case shown has 229 pieces covering 2,190 cells). Decide: a placement of every piece inside the smallest possible square, rotations and reflections allowed. Score: filled fraction of the square × 100, averaged over the grader's cases; C++, 2 s per case.

Four programs of the run on the same case (229 pieces, lower bound 47), each drawn fully placed; S is the side of the square the program found. Data and rendering from coral_tree_report.html.

100 evals, 10 framings scored, 17 records. Best 91.78 at eval 74, inside the framing “cell-first + reinsertion” (opened at eval 38, 47 evals). 3 crashed evals; 10 evals outside the shown range. Run: as embedded in docs/coral_tree_report.html. single-agent and four-agents-on-git values as given in that report; best human as printed.

9 · VLIW kernel scheduling · cycles, lower is better · 567 evals

Given: a tree-walk kernel (hashing and gathers over a forest) and a simulated VLIW machine with 12 ALU, 6 vector, 2 load, 2 store and 1 flow slot per cycle; the output must match bit for bit. Decide: how the work is split between ALU and vector units and how the instructions are scheduled. Score: cycles of the schedule on the grader's forest (lower is better).

Four programs of the run on the grader's forest: each strip is one whole program left to right in time, the five bands are the engines and the filled height of a band is how many of its slots that cycle used. Cycle counts are from the report's re-runs of the rebuilt programs (the run's eval 434 scored 1,097). Data and rendering from coral_tree_report.html.

567 evals, 1 framings scored, 51 records. Best 1,097 at eval 434, inside the framing “one program, refined” (opened at eval 1, 567 evals). 72 crashed evals; 49 evals outside the shown range. Run: as embedded in docs/coral_tree_report.html. single-agent and four-agents-on-git values as given in that report; previous best as printed.