Coral AI Labs
TruffleHog task · coordination illustration

Coral Code

Better work.
With your
coding agent.

We give Codex and Claude Code a network of agents responsible for your code, its connections and the work to be done.

Follow a real Q&A task from TruffleHog, a tool that finds exposed secrets in code.

Scroll to follow the investigation

Your codebase.
Hundreds of agents.

We build a network of file-level agents across your codebase. Each has its own context and responsibility for its assigned code.

File dependencies become communication paths.

On large projects, the network can grow to hundreds of agents as connected files are brought in. Each file has an owner; directional dependencies connect those owners. The network is broader than the agents working at any one moment.

Ask the owners.
Follow the evidence.

We connect your question to the agents responsible for the relevant code. Here, the answer spans detector setup, result filters and reporting.

Relevant owners investigate. Others wait.

The illustration follows file owners responsible for detector setup, result filters and reporting. Other file owners wait for relevant work.

A correct clue.
An incomplete answer.

The baseline reproduced a detector-configuration mismatch. But it did not explain how matches become reported findings.

Four missed checks concerned the notifier and result-consumption path.

The animated exchanges illustrate how file owners share evidence across the result-processing path.

Explain the whole
result path.

Which detectors ran is only the start. Filtering, deduplication and counting determine what the application reports.

The same detector execution can produce different reported counts.

A complete answer explains the notifier’s role in deduplication and metric updates, including what happens when an API caller consumes ResultsChan directly.

From a partial answer
to full marks.

We used Coral Code’s targeted question workflow to score 26/26 on the TruffleHog task, passing all answer checks.

17/26 without Coral Code. 26/26 with Coral Code.

Our research, applied

Beyond retrieving context.
Coordinate the work.

Coral AI LabsWe test how agents work together.
CoralOSWe build shared infrastructure for agents to communicate and coordinate.
Coral CodeWe apply it to your codebase: file-level responsibility, direct communication and controlled changes.

A context tool supplies information to your agent. We also distribute the investigation: agents own their part, exchange evidence, and contribute changes you can inspect.

Latest evaluation · 25 September 2026

43.5%more questions
fully resolved.

99 answers met every check.
30 more than our strongest baseline.

124 SWE Atlas codebase questions
Company-run evaluation · 11 repositories
Opus 5.5 · xHigh
69 of 124 fully resolved
55.6%
Coral Code
99 of 124 fully resolved
79.8%

A fully resolved answer meets every graded rubric item. Highest full-pass rate among seven configurations in our evaluation. Zero-based 0–100% scale.

We used Codex with DeepSeek V4.1 Flash for the outer agent and graph, and Ling 3.0 Flash for initial file agents. The model mix and coordination both changed in this comparison.

Claude Opus 4.5 graded 99/124 versus 69/124. GPT-5.5 rescored the same answers at 96/124 versus 63/124. These are code-understanding results, not completed code changes or measured developer-time savings.

Evaluation details and data ↓

What this can mean for your team

Less time
following up.

A more complete answer can spare an engineer another investigation. That is the economic opportunity, even when the AI costs more.

In this comparison, recorded median inference cost was $3.66 with Coral and $2.00 for Opus. The case for savings rests on the engineering time an additional complete answer avoids.

Illustrative time-value saving

$964per 100 questions

Additional fully resolved answers
24.2
Assumed follow-up avoided per answer
30 min
Assumed loaded engineering cost
$93.40/hr
Modelled extra inference cost
−$166

4.4 minutes of avoided follow-up per additional complete answer covers the modelled extra inference cost.

An illustration, not measured customer savings or a promise of lower payroll. Uses the exact 99/124 and 69/124 rates, with recorded median cost multiplied by volume. Excludes Coral fees, integration, infrastructure, routine review and retries.

Cost basis: selected-attempt API-equivalent prices, excluding graders. DeepSeek uses assumed off-peak pricing; some Coral cost records are incomplete. This illustration uses per-question medians, not the run totals below. At twice the recorded Coral cost, the same illustration yields $598 and breaks even at 14.1 minutes.

What our baseline runs cost

Total inference spend from our reproduction runs, separate from the median-based illustration above.

Astra High
$200
Astra Ultra
$551
Opus 5.5
$351

Our run totals, not Scale AI’s costs. The official leaderboard does not publish inference spend. Unreported means unknown, not $0.

Earlier comparison · context tools

67/67checks passed.

We selected five tasks a single Codex agent had failed. With Coral Code, it passed every check in that sample.

Codex · DeepSeek V4 Pro

Codex alone
45 / 67
With Augment MCP
46 / 67
With Graphify
56 / 67
With Coral Code
67 / 67

Five selected, previously failed tasks. This comparison covers that sample.

Read the results in context

25 September evaluation. We evaluated seven configurations on Scale AI’s SWE Atlas Codebase QnA benchmark. The primary scores use Claude Opus 4.5 as grader. Both headline configurations have grades for all 124 questions. Missing grades in other configurations are excluded from their denominators; refusal scores remain included. These are our results, not a Coral entry on Scale’s public leaderboard.

A second check. On the 120 questions graded for all seven configurations, Coral fully resolved 95 and the strongest baseline resolved 68. A second grader scored the same answer set; it was not a separate replication.

Separate experiments. The TruffleHog story and five-task context-tool comparison are earlier examples, not a trace of the September 25 evaluation. The animated exchanges illustrate the coordination mechanism. Our AgentRadio study tests coordination configurations with the same model; its figures remain separate from these Coral Code results.

Download evaluation totals and savings assumptions (JSON) ↓Product architecture and research at CoralOS ↗

For engineering teams

Measure the value
in your codebase.

Use Coral Code with your coding agent. In team pilots, we will measure investigation time, review and correction time, accepted work and total cost. Those measurements turn benchmark gains into a buying decision.

Explore Coral Code
Continue the researchCoral AutoresearchAgents that learn from experiments ↗