Latest: Coral Code · 25 September 2026
We fully resolved 99 of 124 questions (79.8%), compared with 69 of 124 (55.6%) for the strongest baseline in our seven-configuration evaluation. Every graded check had to pass. Claude Opus 4.5 graded the answers; GPT-5.5 rescored the same answers at 96/124 versus 63/124.
We used Codex with DeepSeek V4.1 Flash and Ling 3.0 Flash. The comparison changes both model mix and coordination. These are our results on Scale AI’s SWE Atlas Codebase QnA benchmark, not an official Coral leaderboard entry. Of the same 124 questions, 62 passed in both configurations, 37 only with Coral, 7 only with the baseline, and 18 in neither.
Read the full evaluation →Download evaluation totals and cost assumptions (JSON) ↓
Earlier research: AgentRadio
We tested 124 SWE-Atlas codebase Q&A tasks across 11 codebases and four languages. Task success required passing all checks.
ConfigurationOpus 4.6DeepSeek V4 Pro
One agent32.3%29.0%
Four agents, divide work39.5%31.4%
+ Negotiation and review51.6%39.5%
+ Live awareness62.1%50.8%
All configurations used Claude Code. The Opus full configuration cost $19.45 per task; one agent cost $2.96. Best-of-six independent attempts scored 37.9% at $17.76. We compared these coordination configurations using Claude Code, with near-matched spending in the DeepSeek comparison.
Read the AgentRadio paper ↗
AgentRadio is an earlier, separate experiment. The opening grid shows the 25 September Coral Code evaluation, with questions grouped by outcome.