Coral AI Labs

Coral Code · our latest evaluation

99/124Codebase questions
fully resolved.

More complete answers can mean less engineering follow-up.

79.8% success · SWE Atlas · 25 September 2026
Company-run evaluation. Every graded check passed.

Scroll to see the engineering value

Illustrative time-value saving

$964per 100 questions.

Assume each additional complete answer avoids 30 minutes of follow-up at $93.40 per engineering hour.

Extra inference cost is already deducted. Modelled value, not measured customer savings. Excludes Coral fees, integration, infrastructure, routine review and retries.

See the cost assumptions ↓

43.5%more questions
fully resolved.

69Opus 5.5 · xHigh
99Coral Code

The same 124 questions. Our strongest baseline versus Coral Code with DeepSeek V4.1 Flash and Ling 3.0 Flash. Model mix and coordination both changed.

Give the codebase
a team of its own.

Coral Code can span hundreds of file-level agents. Relevant owners investigate, exchange evidence and report back to your coding agent.

The diagram shows representative responsibilities, not an agent limit or a replay of the benchmark.

Follow the Coral Code journey →

Minutes matter
more than tokens.

Coral’s recorded median inference cost was $1.66 higher per question. In this illustration, avoiding 4.4 minutes of follow-up per additional complete answer covers that difference.

Median-based illustration per 100 questions · inference plus assumed follow-up
Opus 5.5 · xHigh$2,271

$200 inference + $2,071 assumed follow-up

Coral Code$1,308

$366 inference + $942 assumed follow-up

Illustrative difference$964

12.1 hours of assumed follow-up avoided, less $166 extra inference.

Assumptions: 30 minutes of follow-up for each answer missing one or more checks; $93.40 per loaded engineering hour. Uses exact 69/124 and 99/124 pass rates. Dollar totals are rounded independently. Not measured customer savings or reduced payroll.

Inference uses recorded medians of $2.00 and $3.66 at API-equivalent prices, excluding graders. DeepSeek uses assumed off-peak pricing; some Coral cost records are incomplete. Median × volume is not billed spend. Excludes Coral fees, integration, infrastructure, routine review and retries. At twice Coral’s recorded cost, the illustrative saving is $598.

Explore the economics and evaluation details →

What our baseline runs cost

Total inference spend from our reproduction runs, separate from the median-based illustration above.

Astra High
$200
Astra Ultra
$551
Opus 5.5
$351

Our run totals, not Scale AI’s costs. The official leaderboard does not publish inference spend. Unreported means unknown, not $0.

Your choice of agents.
Our coordination.

We build open infrastructure that lets agents communicate, combine their skills and coordinate work. You choose what to build and how to run it.

CoralOSShared coordination infrastructure
Agents built your way

Any framework. Any language.

Separate contexts. Shared evidence.

Assign responsibility

Give each agent a focused part of the work.

Share discoveries

Keep relevant agents informed as work changes.

Resolve conflicts

Exchange feedback and reconcile competing answers.

Your models. Your tools. Your deployment.

The evidence.

Our latest results, followed by the research behind our approach.

Latest: Coral Code · 25 September 2026

We fully resolved 99 of 124 questions (79.8%), compared with 69 of 124 (55.6%) for the strongest baseline in our seven-configuration evaluation. Every graded check had to pass. Claude Opus 4.5 graded the answers; GPT-5.5 rescored the same answers at 96/124 versus 63/124.

We used Codex with DeepSeek V4.1 Flash and Ling 3.0 Flash. The comparison changes both model mix and coordination. These are our results on Scale AI’s SWE Atlas Codebase QnA benchmark, not an official Coral leaderboard entry. Of the same 124 questions, 62 passed in both configurations, 37 only with Coral, 7 only with the baseline, and 18 in neither.

Read the full evaluation →Download evaluation totals and cost assumptions (JSON) ↓

Earlier research: AgentRadio

We tested 124 SWE-Atlas codebase Q&A tasks across 11 codebases and four languages. Task success required passing all checks.

ConfigurationOpus 4.6DeepSeek V4 Pro
One agent32.3%29.0%
Four agents, divide work39.5%31.4%
+ Negotiation and review51.6%39.5%
+ Live awareness62.1%50.8%

All configurations used Claude Code. The Opus full configuration cost $19.45 per task; one agent cost $2.96. Best-of-six independent attempts scored 37.9% at $17.76. We compared these coordination configurations using Claude Code, with near-matched spending in the DeepSeek comparison.

Read the AgentRadio paper ↗

AgentRadio is an earlier, separate experiment. The opening grid shows the 25 September Coral Code evaluation, with questions grouped by outcome.

Coral AI Labs · Research and applied evaluations

Follow the work.